Did developers opt out of METR's AI study because of selection bias?
METR's August 2025 developer productivity study may have missed its most AI-dependent workers. Understanding whether selection effects—developers refusing to work without AI tools—distorted the speedup estimates matters for interpreting what the data actually shows about AI's real-world impact.
METR's February 2026 update says its second developer productivity experiment no longer gives a reliable reading of how much AI tools speed up experienced open-source work. The study began in August 2025 with 57 developers, and each task was assigned to an "AI allowed" or "AI disallowed" condition. The raw estimates were a speedup of -18% for the 10 original developers who rejoined (confidence interval -38% to +9%) and -4% for the 47 newly recruited developers (confidence interval -15% to +9%). The earlier study had found AI use made tasks "19% longer," with a confidence interval of +2% to +39%. METR believes developers are "likely" more sped up now than in early 2025, but calls its own data "only very weak evidence" for the size of that increase, and both new intervals include zero.
The mechanism METR gives is selection rather than sampling noise. Agentic tools such as Claude Code and Codex made some developers unwilling to take part: "an increased share of developers say they would not want to do 50% of their work without AI," even though the study paid $50 an hour, down from $150 in the original study, which METR thinks "also likely contributed." Task selection narrowed as well. Between 30% and 50% of surveyed developers said they had held back tasks they did not want to do without AI, so the study misses "tasks which have high expected uplift." METR lists further problems it judges smaller: developers running several agents at once found per-task time hard to record, output quality differed between conditions, and one developer completed none of the AI-disallowed tasks assigned. METR reads the selection as pushing its central estimate down, so the figure is "a lower-bound," yet the same passage calls it "likely a bad proxy" for the real productivity effect.
The two nearest notes mark the contrast. The Does Figma Make speed up design task completion? trial reports roughly 20 percent shorter completion times under randomization, and METR's excerpt shows what randomization cannot do alone: it balances only the people and tasks that stay in the study, so a clean design can still miss the users most likely to gain. The time-logging problem is the measurement side of the reallocation described in Does AI really save time, or just change how we spend it?, where time moves toward prompting and evaluating output. METR's developers found that shift makes time records unreliable when they worked an unrelated task while an agent ran. METR's planned use of observational data, such as aggregate commit statistics and transcripts, points toward the trace-based approach in Does generative AI shift knowledge workers away from communication?, which avoids self-reported time and task choice but measures activity rather than output.
The excerpt does not establish how large AI's effect on these developers is, or whether it has grown since 2025. METR's growth claim rests on conversations and surveys, and it calls the data "very weak evidence" for that claim. It also says developer self-reports of "very high speedups" "can be quite unreliable," which leaves the study without a usable estimate of either size or direction. The sample is experienced open-source contributors with a median of 10 years' experience, working on their own repositories across 143 repos and 800+ tasks, so it does not describe developers in other settings. The strength the evidence supports is narrow: this randomized design has stopped measuring what it was built to measure, and METR is redesigning it. The implication is that any claim about the size of the speedup needs a design that keeps high-adoption developers in the sample. METR's proposed options, including more intensive experiments, fixed-task experiments and developer-level randomization, aim at that, but the excerpt reports none of them as results.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What human oversight must AI research systems have?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does Figma Make speed up design task completion?
Does access to a prompt-to-design tool reduce the time needed to complete structured design work, and does the effect differ between professional designers and product managers?
contrasts: a randomized speedup estimate, where METR's excerpt shows randomization cannot fix who stays in the study
-
Does AI really save time, or just change how we spend it?
Explores whether AI's time savings are real or illusory—whether the time freed from direct work simply shifts to AI interaction tasks like prompt composition and output evaluation, with different cognitive and learning consequences.
extends it to measurement: developers working unrelated tasks while agents ran made time-per-task logs unreliable
-
Does generative AI shift knowledge workers away from communication?
When knowledge workers adopt generative AI heavily, do they spend proportionally more time on individual documentation and less on coordination with colleagues? Understanding this matters because it suggests AI may reshape not just productivity but the social fabric of how teams work together.
parallel observational route: METR plans commit and transcript data, which trace studies of office work illustrate
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- We are Changing our Developer Productivity Experiment Design
- How much does AI impact development speed? An enterprise-based randomized controlled trial
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- How AI Impacts Skill Formation
- How AI Can Degrade Human Performance in High-Stakes Settings
- Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI
- Microsoft New Future of Work Report 2025
- Generative AI at Work
Original note title
METR says selection effects make its August 2025 developer study an unreliable signal of AI speedup — developers opted out rather than work without AI