Developers Using AI Were 19% Slower. They Reported Being 20% Faster.

METR ran a randomized controlled trial on 16 experienced open-source developers across 246 real tasks, in repositories they had worked in for an average of five years. Half the tasks allowed AI tools, half did not.
The developers predicted AI would make them 24% faster. They were 19% slower on the tasks where they used it. Then, after finishing, having personally lived through the slowdown, they estimated AI had made them 20% faster.
That is a 39-point gap between what was measured and what was felt, and it survived direct experience of the opposite.

Hold that number next to the surveys. Deloitte reports 84% of companies say they are gaining ROI from AI. The 2025 DORA report finds over 80% of developers believe AI increased their productivity. Those are real findings, honestly collected. They are also measurements of the same feeling METR caught being wrong.
Almost nobody has the instrument
McKinsey’s State of AI survey found that fewer than one in five organizations track well-defined KPIs for their gen-AI solutions. The same survey found KPI tracking sits among the practices with the most impact on EBIT. The biggest lever in the study is the one four out of five companies skip.
The result shows up at the top. PwC’s 29th Global CEO Survey, published January 2026 and covering 4,454 chief executives across 95 countries, found 56% report neither higher revenue nor lower costs from AI over the past year. Only 12% report both. Those 12% are two to three times more likely to have put AI into products, demand generation, and strategic decisions rather than running isolated pilots.
Meanwhile S&P Global Market Intelligence found 42% of companies abandoned most of their AI initiatives in 2025, up from 17% the year before, with the average organization scrapping 46% of proofs-of-concept before production.
Read those together and a specific failure appears. It is not that companies tried AI and it did not work. It is that most of them cannot tell you either way, so the decision to continue or abandon gets made on impressions. And we know what impressions are worth here: 39 points.
What “adopted” would have to mean
The only large-scale measurement that ties AI adoption to outcomes rather than opinions is DORA, which surveyed nearly 5,000 technology professionals in 2025. Its findings are worth reporting in full, including the part that reversed.
In 2024, a 25% increase in AI adoption was associated with an estimated 1.5% decrease in delivery throughput and a 7.2% decrease in delivery stability. In 2025, the throughput relationship flipped positive. The stability relationship did not. It stayed negative both years.

We are not going to smooth that over. Two consecutive waves of the same instrument disagree on the direction of the throughput effect, and no one has explained why. The plausible readings: tooling improved between waves, the adopting population changed as usage went from 76% to 90%, or the survey is capturing which teams adopt rather than what adoption does. DORA’s own framing points at the third and is the most useful sentence in the report: AI does not fix a team, it amplifies what is already there.
The stable finding is the one nobody quotes. AI adoption degrades delivery stability, in both waves, without exception. Which gives a definition worth writing down:
A workflow is adopted when you can state its pre-AI baseline, the evaluation it has to pass, the one outcome number it moved, and what it did to your failure rate.
Throughput without the failure rate is half a measurement, and it is the half that flipped.
The instrument is an eval, and it is smaller than you think
Anthropic’s guidance on evaluating agents is the most practical writing available on this, and the scale it recommends is modest: 20 to 50 tasks drawn from real failures is a good start. Not thousands. Convert the manual checks you already run before each release into tests.
The single sharpest test for whether an eval is real: a good task is one where two domain experts would independently reach the same pass or fail verdict. If your team cannot agree on what counts as a correct answer, you do not have a measurement problem, you have a specification problem, and no dashboard will rescue it.
Three rules that keep evals honest. Grade the output, not the path, because demanding a specific sequence of steps punishes valid solutions. Test both the cases where a behavior should happen and where it should not, since one-sided evals produce one-sided optimization. And treat a 0% pass rate across many trials as a broken task rather than an incapable system.
The operator payoff is the part that gets missed. Teams with evals can assess a new model, tune, and upgrade in days. Teams without face weeks of testing every time the ground moves. The eval is not overhead on top of the work. It is the thing that lets you swap the model underneath the work.
Knowing when to stop
Every framework for deciding whether to scale, pivot, or kill an AI pilot is invented. Four-to-six week windows, phase gates, fourteen-element pilot charters: none of it has outcome data behind it. We looked.
Two findings from outside AI do the job better.
The first is a base rate. In Ronny Kohavi’s experimentation work at Microsoft, of well-designed experiments intended to improve a key metric, about a third succeeded, a third were flat, and a third were significantly negative. That was after heavy upfront scoping. The median organization sees roughly 10% of its experiments move the number they targeted. So a failed AI pilot is not a verdict on AI. It is the base rate for changing anything.
The second explains why the failures do not get cut. Barry Staw’s 1976 study “Knee-deep in the Big Muddy” put 240 subjects through a simulated investment decision and found that people committed the most additional resources to a failing course of action precisely when they were personally responsible for the original decision. Fifty years later that is still the mechanism killing AI budgets.
The prescription follows from the research rather than from a framework: the person who championed the build must not be the person who decides whether it lives. Set the success criterion before the work starts, and hand the kill decision to someone with no authorship in it.
Most companies do not have an AI problem. They have an instrumentation problem wearing an AI costume. The 12% that are getting both revenue and cost gains did not find a better model. They built something that could tell them whether it was working.
Was this useful?
Get the next one
New pieces when they land. No cadence promises, no noise.