Nobody abandoned your product because it was 3% less accurate

Two numbers, from the same year, describing the same market.

McKinsey reports that 91% of employees say their organisation uses at least one AI tool. Pew's survey of US workers found 21% actually use AI at work, with daily use at around 10%.

Both are true. They measure different things, and the gap between them is where most AI products quietly die.

More precisely: BCG's research found that over 85% of employees remain at early stages of AI adoption, using tools mainly for basic information retrieval, and fewer than 10% reach the point where AI is central to how they do their core work. S&P Global reports that 42% of companies have abandoned most of their AI projects. And the abandonment has a characteristic shape — enthusiasm at launch, decay over the following 30 to 90 days, then a usage report showing a handful of logins against a lot of paid licences.

Almost none of that is a model quality problem. The products largely worked. People stopped using them.

Which is worth stating plainly, because it inverts where most teams spend their effort: for the large majority of AI products, the binding constraint is adoption, not accuracy. Another two points of benchmark performance changes nothing if the tool sits outside the workflow, the user can't correct it, and using it visibly carries a professional risk.

Your adoption metric is probably lying to you

Before the design argument, a measurement one — because most teams don't know they have an adoption problem until it's terminal.

Adoption surveys and internal dashboards suffer from what's fairly called headline inflation: an organisation plans to adopt AI, one person uses ChatGPT, and it's reported as adoption. The same distortion happens inside products. Someone opened the feature once. That's a monthly active user.

The cleanest available public time series illustrates the spread. Gallup's Q2 2026 survey of 22,573 employed US adults found 52% using AI in their role at least a few times a year — but 30% weekly or more, and 15% daily. Three defensible numbers, differing by more than 3×, depending on which threshold you pick. Firm-level counts land near 18%.

The metrics that actually predict whether a product survives:

  • Depth, not breadth. What share of the target workflow runs through it, for the people who use it at all?
  • Return usage after the novelty window. Week 1 usage is curiosity. Week 6 is adoption.
  • Task completion, not session count. Did they finish the job in the tool, or start it there and finish it elsewhere?
  • Rework rate. High usage with heavy downstream correction is worse than low usage. It looks productive on the dashboard and bleeds value — one useful framing calls this the AI tax: if a tool costs $30 per user per month, it has to save each user more than $30 of net productivity after rework, or the investment is underwater.
  • Voluntary versus mandated use. If usage collapses when the mandate lifts, you have compliance, not adoption.

What forty years of research says about trusting machines

Here the literature is unusually clear, and unusually ignored by product teams.

The foundational finding is algorithm aversion (Dietvorst, Simmons and Massey, Journal of Experimental Psychology: General, 2015). Across five studies, people who watched an algorithm make forecasts became less likely to choose it over an inferior human forecaster — and this held even among participants who had seen the algorithm outperform the human. The mechanism: people lose confidence in an algorithm faster than in a human after observing the same mistake.

A human colleague who gets something wrong is having a bad day. A system that gets the same thing wrong is broken. The error is identical; the attribution isn't.

For a product team this is uncomfortable, because it means accuracy improvements are not linearly convertible into trust. You can ship a system that beats the human baseline, demonstrate it, and still lose the user at the first visible error.

But the picture is more actionable than that, because later work found the boundaries.

Aversion is relative, not absolute. MIT Sloan research by Zhang and Gosline found that algorithm aversion disappears as the AI's error gets smaller or the human's gets larger — and, critically, that aversion is triggered not by the AI performing worse than expected in some abstract sense, but by the AI's error exceeding the error the user expected from a human.

That single finding has a direct product consequence: the reference point is a design variable. If your marketing promises flawless performance, you've raised the expected standard above any human benchmark and guaranteed that the first error reads as failure. If you state honestly what the system gets wrong and how often, the same error lands as within expectations. Overclaiming doesn't just risk credibility — it mechanically manufactures the abandonment condition.

And the strongest lever of all is control. Dietvorst, Simmons and Massey followed up in Management Science (2018) with a paper whose title is the whole finding: Overcoming Algorithm Aversion: People Will Use Imperfect Algorithms If They Can (Even Slightly) Modify Them.

Even slightly. Giving users a small amount of adjustment over an imperfect system substantially increases their willingness to use it — not because the modification improves accuracy, but because it restores agency.

Most AI product design does the opposite. The output arrives finished. Accept or reject. That binary is exactly the condition under which a single error becomes abandonment.

What this looks like in practice:

  • Outputs that are editable in place rather than regenerable only
  • Confidence thresholds the user can set, not ones you set for them
  • Override that persists — if I correct it once, it shouldn't make the same call tomorrow
  • Showing the reasoning or the source so the user can check the part they doubt
  • A visible dial between "suggest" and "act" so users can advance at their own pace

None of that improves the model. All of it improves adoption.

The social cost nobody designs for

Here's an adoption barrier that appears in almost no product spec.

Research from BetterUp Labs and Stanford, reported in Harvard Business Review, found that half of surveyed workers view colleagues who send low-quality AI-generated work as less capable, and 42% view them as less trustworthy.

So using your AI tool carries a professional risk that has nothing to do with your product's quality. If the output is visibly machine-made and imperfect, the user — not the vendor — absorbs the reputational damage. A rational employee responds by using the tool privately, editing heavily, or not using it for anything that will carry their name.

This explains a pattern teams find baffling: high usage on low-stakes tasks, near-zero on high-stakes ones. That isn't a capability gap. It's a risk calculation, and it's correct.

Design consequences:

  • Never make output attributable before the user has reviewed it. Draft-first, always.
  • Make editing feel like authorship rather than correction.
  • Don't add visible "generated by AI" markers to internal work products unless disclosure genuinely requires it — you're taxing your own user.
  • Reduce the tells that make output identifiably machine-made and low-effort.

Gallup found only about 1 in 10 employees feels comfortable using AI in their role. Some of that is skill. A meaningful part is this.

Placement beats capability

The most common structural adoption failure is the simplest: the tool lives somewhere the work doesn't.

If your users work in a helpdesk, a CRM, an inbox or a spreadsheet all day, a separate interface — however good — is asking them to add a context switch to every task. Switching costs are paid every single time; the benefit is paid once per task. For anything short of a dramatic improvement, the arithmetic doesn't work, and users revert without ever deciding to.

The corollary is that the same capability delivered in the workflow will out-adopt a better capability delivered adjacent to it. This is why embedded features frequently beat superior standalone products, and it's the single most reliable design decision available.

A related placement question: does the tool arrive before or after the decision? Analysis delivered after someone has committed is filed, not used. Intelligence that reaches a person while they still have a choice to make is the only kind that changes behaviour.

Training is the highest-return intervention, and it's unglamorous

The evidence here is stronger than for almost any product-side intervention.

Research from the London School of Economics with Protiviti found that 93% of employees who received AI training used the tools regularly, against 57% of those without. BCG identified a threshold: employees receiving at least five hours of AI training show significantly higher regular usage and confidence, with in-person coaching making the largest difference.

Five hours. That's the intervention that moves adoption from 57% to 93%, and it costs a fraction of the engineering time typically spent chasing accuracy improvements that, per the research above, won't convert into trust anyway.

For product teams shipping to customers, the equivalents are onboarding depth, worked examples using the customer's own data, and a first-run experience that produces one real completed task rather than a feature tour.

Design for month two

Abandonment clusters at 30 to 90 days. Enthusiasm carries the first few weeks; what happens next determines everything.

Month one is the demo working on the cases people tried first. Month two is the awkward input, the edge case, the wrong answer nobody caught, the moment the user has to decide whether this is a tool that occasionally errs or a tool that can't be trusted. Per the aversion research, that decision is made fast and is hard to reverse.

So build for it deliberately:

  • Ship the boring reliability first. Consistency on common cases beats brilliance on rare ones, because common cases are where trust is formed.
  • Fail visibly and well. A system that says "I'm not confident about this one" preserves trust; one that guesses confidently and is wrong destroys it. Uncertainty communicated is credibility earned.
  • Make correction cheap and consequential. If fixing an error takes longer than doing the task manually, the user has learned the tool costs more than it saves.
  • Watch the second-month cohort separately. Aggregate usage hides decay for months. Cohort it.
  • Expand scope only when the current scope holds. The teams that succeed start narrow, run alongside the human process, and widen once the numbers justify it. That's an adoption method, not caution.

The organisational half

Two findings worth carrying into any rollout conversation.

WRITER's 2026 survey of 1,200 employees and 1,200 executives found 75% of executives admitting their AI strategy is "more for show" than actual internal guidance, and 48% describing adoption as a massive disappointment. Meanwhile 60% reported plans to disadvantage non-adopters. That combination — no real strategy, plus pressure to adopt — produces defensive, minimal, performative usage. People log in enough to be seen logging in.

Adoption driven by fear produces compliance metrics and no value. Adoption driven by a tool that visibly makes someone's day better produces the other thing. Only one of them survives the mandate being lifted.

What to do this week

  1. Re-cut your adoption numbers by depth and by cohort. Weekly-or-more usage, by signup month. If month-two retention is falling, you have your priority.
  2. Find the modification affordance. Where can a user adjust, override or edit your output? If the answer is nowhere, that's the highest-leverage change available and it's usually small.
  3. Check the reference point your marketing set. If you promised flawless, you've guaranteed that the first error reads as failure. Publish honest limitations instead.
  4. Map the context switches. How many does a user pay to complete one task? Every one is a tax charged per use.
  5. Ask what using your tool costs the user socially. Whose name goes on the output, and is it reviewable before it does?
  6. Commit five hours of real training for internal rollouts. It's the best-evidenced intervention in this entire article.

The uncomfortable summary is that the highest-leverage work on most AI products isn't in the model. It's in whether a person can correct it, whether it sits where they already work, whether it makes them look competent, and whether the second month is as good as the first. Those are product decisions, and they're cheaper than the accuracy work most teams are doing instead.

Sources: Dietvorst, Simmons & Massey, "Algorithm Aversion: People Erroneously Avoid Algorithms After Seeing Them Err," Journal of Experimental Psychology: General (2015); Dietvorst, Simmons & Massey, "Overcoming Algorithm Aversion," Management Science (2018); BetterUp Labs & Stanford, via Harvard Business Review; London School of Economics with Protiviti; Gallup; McKinsey; Pew Research; BCG; S&P Global.

Trying to work out why an AI feature that tested well isn't sticking after week one? Get in touch.