AI can write and fix code well now — so why do humans still have to check its work before it ships?
Why haven't AI agents replaced human code review workflows?
This explores why AI agents, which now write and fix code well, haven't taken over the human job of reviewing code. The collection has no studies of code review itself, so this answer draws on nearby research about AI as an evaluator, AI on long tasks and how organizations adopt new tools.
This explores why AI agents, which now write and fix code well, haven't taken over the human job of reviewing code. A note up front: the collection has no studies of code review itself. What it does have is a close cousin: research on whether AI can replace human reviewers of scientific papers, and on how reliable AI is as a judge of other AI's work. That body of work points to a clear answer. Reviewing is a different kind of job from producing. AI is getting much better at producing, but reviewing still runs into problems that better coding skill doesn't fix.
The sharpest evidence comes from peer review. One study found that AI reviewers share a 'hivemind': they agree with each other more than human reviewers do, so adding more AI reviewers adds less independent scrutiny than it seems to Can AI systems safely replace human peer reviewers?. Worse, simply rewriting a paper's wording raised AI scores without improving the science at all. Apply that to code. A review step that can be gamed by surface polish, such as tidy variable names, confident comments or a well-written pull request description, isn't really a check. Human reviewers bring different blind spots, and that variety is much of the value. The AI Scientist project shows the risk from the other side: an AI system wrote a paper, reviewed it with its own panel of AI reviewers, and passed a workshop's first round Can one AI system complete a full research cycle end-to-end?. That is impressive, but it is exactly the closed loop a human reviewer exists to break.
AI evaluation isn't hopeless, though, and the details are telling. Giving the judge agent tools to gather evidence, rather than just reading the output and giving an opinion, made its verdicts far more stable. 'Judge shift' fell from 31% to 0.27% Can agents evaluate AI outputs more reliably than language models?. But the agent's memory module passed its own errors down the chain. In government consultation analysis, AI disagreed with experts about as much as the experts disagreed with each other Does AI theme-mapping perform as well as human reviewers?. Together these suggest AI review can match human review on well-bounded judgments. They also suggest it fails in ways that are correlated and hard to see, while human disagreement tends to be spread out and easier to spot.
Two more threads explain why teams keep humans in the loop. First, time horizon: METR found agents beat expert humans 4× on two-hour tasks, but humans pulled ahead with 8 to 32 hours When do AI agents outperform human research experts?. Good code review often depends on that slow-building knowledge: why the codebase is shaped the way it is, what broke last year, where the system is heading. Second, context. AI works from a shifting mix of prompt, retrieved files and conversation history rather than a stable understanding How does AI context differ from conventional software context?. A reviewer's job is largely to hold a stable picture of the system. The practical response in the research isn't replacement but structure: wrap coding agents in orchestration layers that leave a step-by-step record a human can audit and recover from Can orchestration layers make coding agents more auditable?.
The insight you might not have expected is that self-improving agents already rely on automated review. They just don't call it review. Systems like the Darwin Gödel Machine and its relatives improve by testing each variant against benchmarks and keeping what works Can AI systems improve themselves through trial and error? Can an AI system improve its own search methods automatically?. When 'good' can be checked by running a test, the reviewer can be automated. Code review is mostly about everything a test can't capture: intent, maintainability, fit with the team's direction. Then there's the organizational layer. Even when the technology works, adoption depends on people recognizing the task as automatable and on decisions that cut across teams Does easier tool-building actually solve enterprise adoption problems?. Review is also where a team spreads knowledge and assigns responsibility, and AI doesn't replace either.
Sources 10 notes
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
The AI Scientist performed ideation, coding, experiments, writing, and self-review autonomously, producing a manuscript that passed the first round at a machine learning workshop with 70% acceptance rate. Five ensemble reviewers and an area-chair model judged the output against NeurIPS guidelines.
Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.
UK government's Consult tool achieved F1 0.76 against expert reviewers, compared to F1 0.81 between two human reviewers. Differences rarely affected which themes ranked top, suggesting AI performance was competitive despite inherent subjectivity in theme assignment.
METR's RE-Bench found AI agents score 4× higher than expert humans at 2-hour budgets but humans narrowly exceed agents at 8 hours and lead 2× at 32 hours, suggesting agents hit scaling plateaus while humans improve with extended effort.
Show all 10 sources
AI interactions operate on a substrate of constantly shifting context—prompt, history, retrieved data, hidden state—that users cannot internalize like traditional UIs. This structural mutability demands a new design discipline centered on context engineering rather than interface design.
Dr. Claw wraps existing coding agents in persistent state objects and skill libraries, reporting higher research completeness and a traceable, recoverable process trail while keeping the underlying executor unchanged.
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
An outer loop successfully read inner loop code, identified bottlenecks, and generated new Python mechanisms at runtime, discovering combinatorial optimization and bandit methods that broke the inner loop's deterministic patterns and improved performance on GPT pretraining by 5x.
Evans argues that reducing coding friction masks two structural barriers: most workers don't see their own tasks as automatable, and enterprise adoption requires organizational decisions that span departments and timelines—not just technical capability.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- Stop Automating Peer Review Without Rigorous Evaluation
- Towards End-to-End Automation of AI Research
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search