Skip to content

Developers felt faster. They were slower

A controlled trial found experienced developers were 19% slower with AI while believing it sped them up. A year later the same developers were faster. The gain was never in adoption. It was in aiming AI at the right work.

The most quoted AI study of 2025 found that experienced developers were slower with AI, not faster. It got read as proof the tools are oversold. That is the wrong lesson. The developers who felt fastest were the ones losing the most time, and the fix was not to drop the tools. It was to learn where they actually help.

What you need to know

  • In a randomised controlled trial, experienced open-source developers took 19% longer to finish tasks when allowed to use AI, on their own repositories.
  • They did not notice. They expected a 24% speed-up before, and still believed AI had sped them up by 20% after. The perception gap is the finding, not a footnote.
  • A year later, METR re-ran it with many of the same developers. The slowdown had flipped to a speed-up. Nothing changed about the tools' potential. What changed was skill at pointing them at the right work.
  • The lesson for leaders: the return on AI is not in getting people to use it. It is in getting people to use it on the tasks where it wins, and to stop where it loses. That is a focus problem, not an adoption problem.

19%

longer to complete tasks when experienced developers were allowed AI, in a controlled trial on their own repos

Source: METR, July 2025

20%

speed-up those same developers believed they had gained, while measurably losing time

Source: METR, July 2025

16

experienced open-source developers, across 246 real tasks on repositories they knew well

Source: METR, July 2025

What the study actually found

METR ran a randomised controlled trial with 16 experienced open-source developers across 246 real tasks on repositories they maintained, some with over a million lines of code. Half the tasks allowed AI, mostly Cursor Pro with Claude 3.5 and 3.7 Sonnet. The other half did not. The result surprised everyone, including the developers: the AI-allowed tasks took 19% longer.

The number that should stop a board is not the 19%. It is the gap between what happened and what people felt. Before the trial, the developers expected AI to speed them up by 24%. After it, having been measurably slower, they still believed it had sped them up by 20%. They were wrong about their own productivity by roughly 40 percentage points, on work they do every day.

This is the same shape as the AI hangover: AI changes how work feels long before, and sometimes instead of, changing what work produces. Feeling faster and being faster are different measurements, and only one of them shows up in your delivery.

Why they slowed down

The developers were not fumbling. They were experts on code they knew cold. The time went into a tax that is easy to miss: prompting, waiting, reading a generated diff, deciding whether to trust it, and reworking the parts that were subtly wrong. On a codebase where you already know exactly what to type, that loop costs more than it saves. AI was strongest on the unfamiliar and the boilerplate, and weakest precisely where these developers were strongest.

That is the whole point, and it is why "are we using AI?" is the wrong question. AI is not uniformly fast or slow. It is fast on some tasks and slow on others, and the difference is large enough to flip a team from a gain to a loss. A rollout that treats adoption as the goal drives usage up without asking the only question that moves delivery: used on what?

The follow-up nobody quotes

Here is the part that got buried. In early 2026 METR revisited the question with many of the same developers. The slowdown had reversed: the central estimate landed on a net speed-up of around 18% for that group, though METR is careful to call it weak evidence, because the developers keenest on AI increasingly refused to work without it, which skews who stays in the sample.

Take the caveat seriously and the direction still holds: the same people, on the same kind of work, went from slower to faster inside a year. The models improved, but the bigger shift was human. They learned which tasks to hand over and which to keep, when to accept a diff and when to close the tab. They got good at aiming.

The first study did not show that AI makes developers slower. It showed that AI makes untuned developers slower, and that untuned developers cannot feel it. The gap between those two readings is the difference between abandoning the tools and learning to aim them.

Mak KhanChief AI Officer

What this means for anyone rolling out AI

The same trap sits under every AI programme, not just engineering. Give a team a capable tool, watch usage climb, and assume value follows. It does not follow on its own. Value shows up when people learn the tool's edges: where it wins, where it loses, and how to tell the difference on the task in front of them. That learning is the work, and it is the part most rollouts skip.

Three things this changes:

Aim before you scale. Pick the tasks where AI clearly wins for your team and start there, rather than pushing it across everything and hoping. The wrong first target teaches people the tool does not work, and that lesson is expensive to unteach. This is what an AI roadmap is for: deciding where AI earns its place before you spend a year finding out the hard way.

Trust the numbers over the feeling. The developers were confident and wrong. Measure delivery on the actual tasks, before and after. If you cannot show the gain in output, the reported time-savings are a feeling, not a result.

Budget for the learning curve, not just the licences. The reversal took the same people a year of doing the work. A three-month pilot that measures a slowdown and cancels is measuring the cost of learning and calling it the cost of the tool.

None of this is an argument against AI. It is the case for building the right thing: the study that looks like a warning about AI is really a warning about focus. The teams that win are not the ones that adopt fastest. They are the ones that learn fastest where AI actually helps, and have the discipline to stop where it does not.