Do AI coding tools actually speed up experienced developers?
Developers predicted AI tools would make them 24% faster, but a randomized trial measuring real work found the opposite. Understanding this gap between forecast and outcome matters for assessing AI's real productivity impact.
In a randomized controlled trial, allowing AI tools made experienced open-source developers slower: completion time rose 19 percent. Sixteen developers completed 246 tasks, each randomly assigned to allow or disallow early-2025 AI tools, mostly Cursor Pro with Claude 3.5/3.7 Sonnet. Before the tasks they forecast a 24 percent reduction in completion time, and afterward they estimated a 20 percent reduction. The 19 percent is the trial's measurement. The forecasts and estimates are the developers' own, along with economics and ML experts' forecasts of 39 and 38 percent reductions, and the excerpt's authors say both groups "drastically overestimate the usefulness of AI on developer productivity, even after they have spent many hours using the tools."
The design runs on real work. Each developer lists real issues from repositories they regularly contribute to, which average 23,000 stars and 1,100,000 lines of code. Issues are randomized by a simulated coin flip, and the outcome is time to completion, not lines of code or pull requests. The authors note that earlier field studies used those outcomes, though AI can move them without productivity rising. To explain the result, the authors hand-label 143 hours of screen recordings (29 percent of developers' total hours), pull source-code statistics, and survey and interview participants. Of 21 candidate factors, they find evidence that 5 contribute, mixed or no evidence for 10, and evidence against 6. The likely contributors include over-optimism about AI usefulness, high developer familiarity with the repositories, large and complex codebases, low AI reliability (developers accept under 44 percent of generations and spend 9 percent of their time reviewing and cleaning AI output), and implicit repository context.
The result cuts against the speedup figures elsewhere in the library. A randomized trial of 50 designers and 50 product managers links Does Figma Make speed up design task completion? to shorter completion times, so the two trials measure similar outcomes and point in opposite directions. The excerpt's own account of where AI helps is narrow: the slowdown is tied to mature repositories and developers' familiarity with them, and the authors expect that "small greenfield projects or development in unfamiliar codebases" may see substantial speedup. Anthropic's cited speedup of about 3x to about 52x, discussed in Is AI development already being handed to AI systems?, is a different kind of evidence. The excerpt argues for "field experiments with robust outcome measures, compared to relying solely on expert forecasts or developer surveys," and that standard does not reach the lab's own figure. The excerpt also points at where the time goes, decomposing it at about 10-second resolution, but it does not report that decomposition, so it cannot confirm Does AI really save time, or just change how we spend it?.
The excerpt does not establish a general productivity loss. The authors say the slowdown "does not imply that current AI tools do not often improve developer's productivity," and they allow that future models, better scaffolding or domain fine-tuning could change the sign. The evidence is 16 developers in one setting, and the 21 factors are hypotheses, not tested causes. The authors concede that experimental artifacts cannot be entirely ruled out, though they find the slowdown robust across analyses. The implication, at the strength the evidence allows: for experienced developers on mature codebases they know well, early-2025 tools slowed completion, and developers' own speedup estimates are weak evidence about that setting. Claims beyond it need their own measurement.
Inquiring lines that read this note 25
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do AI coding tools measurably improve developer productivity and code quality?- Do AI coding tools improve code quality alongside task speed?
- Why do experienced developers benefit more from AI coding assistance?
- How much time do developers spend verifying AI-generated code?
- Does AI coding assistance help junior developers close skill gaps?
- Does high-level design work benefit differently from AI than routine coding tasks?
- Why do developer self-reports of AI speedups tend to be unreliable?
- How much time do developers spend reviewing and fixing AI code?
- Does AI help more on small greenfield projects than mature codebases?
- Why did developers and experts forecast such large AI productivity gains?
- Can AI design tools like Figma Make show speedups when coding tools show slowdowns?
- Why do novice engineers lose confidence in coding after using AI tools?
- Does prior IDE tool use predict stickiness with new coding assistants?
- Why did programmer headcount not shrink after AI coding tools arrived?
- Why do experienced developers report slower task completion with AI assistance?
- Does AI-assisted coding actually speed up experienced developers?
- How much can self-reported AI use tell us about actual productivity changes?
- How do time-logging problems distort AI productivity measurement in developer studies?
- Why do most organizations lack reliable data on AI's actual impact on productivity?
- How much does AI actually automate versus augment in real workplace tasks?
- Does checking AI output carefully eat back most of the time it saves?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does Figma Make speed up design task completion?
Does access to a prompt-to-design tool reduce the time needed to complete structured design work, and does the effect differ between professional designers and product managers?
contrast: randomized trial with the opposite sign, on structured design work rather than mature repositories
-
Is AI development already being handed to AI systems?
Anthropic reports rising task length, code authorship, and speedup metrics as evidence that AI systems are taking on development work. The question is whether these measures actually demonstrate autonomous delegation of R&D or reflect improvements in assisted productivity.
contrast: Anthropic's own speedup figure against an outside, measured completion-time outcome
-
Does AI really save time, or just change how we spend it?
Explores whether AI's time savings are real or illusory—whether the time freed from direct work simply shifts to AI interaction tasks like prompt composition and output evaluation, with different cognitive and learning consequences.
same time-on-task question; this excerpt names the method but does not report its time breakdown
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- We are Changing our Developer Productivity Experiment Design
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- How much does AI impact development speed? An enterprise-based randomized controlled trial
- How AI Can Degrade Human Performance in High-Stakes Settings
- How AI Impacts Skill Formation
- What 81,000 people told us about the economics of AI
- 2025 Stack Overflow Developer Survey: developers remain willing but reluctant to use AI
- Beyond Productivity: Measuring the Real Value of AI
Original note title
a randomized trial found early-2025 AI tools slowed experienced open-source developers by 19 percent — developers had forecast a 24 percent speedup