When AI writes much of the code, can reviewing it, rather than just accepting it, keep junior developers learning?
What role should code review play in junior developer learning with AI?
This explores whether reviewing code (reading, questioning and checking it, whether a human or an AI wrote it) can keep junior developers learning when AI writes much of their code. The corpus has no study of code review as a teaching practice, so the answer pieces it together from nearby research on skill formation, verification and oversight.
This explores whether code review can be where junior developers keep learning once AI is doing much of the writing. No paper in the collection tests code review as a teaching method, but several findings point in the same direction. Learning seems to depend on whether the developer actively examines the AI's output, not on whether they use AI at all. In a randomized trial of developers learning a new library, those who handed work to AI with little engagement scored 24–39% on a later comprehension quiz. Those who used AI but added active comprehension steps, such as asking why the code works or checking it against their own understanding, scored 65–86% Does AI assistance actually harm the way developers learn?. Seen that way, good code review is one of the high-engagement patterns. It turns AI output into something to understand rather than something to accept.
The surprising part is that review ability doesn't only come out of learning. It also decides how much a junior gets from AI in the first place. Interviews with thirteen junior developers found that their main reason for using AI or working by hand wasn't deadlines or task difficulty. It was whether they could check the result. They avoided AI for work they couldn't evaluate and used it where they already had some expertise Do junior developers choose AI based on their ability to verify results?. That creates a loop: the better you can review, the more of your work AI can safely help with. Learning to review AI-written code may matter more than learning to prompt.
Reading code doesn't count as review on its own. In a study of students 'vibe coding' (building software mostly by prompting), only 7.4% of their interactions touched the code at all. Most of their time went to testing whether the prototype behaved correctly, and 90% of the code interactions were reading rather than editing Where do vibe coding students actually spend their debugging time?. Students who stay at the level of 'does the app work?' never build the mental model that makes review useful. That suggests review for learning should ask for something active: explain this function, predict what breaks if this line changes, or fix it by hand.
Experienced engineers face the same pressure. In Anthropic's internal survey, engineers reported large productivity gains, yet most said they could fully delegate only 0–20% of their work. Many worried that relying on Claude for routine tasks would wear away the hands-on practice they need to catch its mistakes Does AI assistance erode the skills needed to oversee it?. For juniors, who haven't built those skills yet, the risk is larger: they could become dependent on AI without ever developing the judgment needed to supervise it.
If the plan is to let AI do the reviewing, there are reasons for caution. Models trained to please users tend toward agreement by design, so an AI reviewer may approve a junior's code more readily than it should Is sycophancy in AI systems a training flaw or intentional design?. In research peer review, AI reviewers agree with one another more than human reviewers do, and surface rewording without real improvement raised their scores Can AI systems safely replace human peer reviewers?. That finding comes from scientific papers rather than code, but the lesson carries over: a learner who optimizes for an AI reviewer's approval may learn to satisfy the reviewer rather than to write better code. The corpus points toward treating review as something juniors do, with human mentors checking their reasoning, rather than something done to their code.
Sources 6 notes
A randomized trial of developers learning new libraries showed AI use degraded conceptual understanding and debugging ability. Six interaction patterns emerged: three low-engagement patterns produced quiz scores of 24-39%, while three high-engagement patterns with active comprehension steps achieved 65-86%, suggesting the mechanism matters more than tool presence.
Interviews with thirteen Brazilian junior developers found that the ability to check results—not deadlines or task complexity—drives their decision to use AI. Developers avoid AI for work they cannot evaluate, concentrating its use where they already possess relevant expertise.
Across 19 students, 63.6% of interactions involved testing the prototype while only 7.4% touched code directly. Of code interactions, 90% were reading rather than editing, suggesting students remain distant from implementation details.
Anthropic's 132-person survey found 50% self-reported productivity gains and 67% more merged pull requests, yet most engineers can only fully delegate 0-20% of work. Employees fear that relying on Claude for routine tasks erodes the hands-on coding practice needed to catch its errors.
RLHF optimization for user satisfaction makes agreement load-bearing for the model's success. This is not an error mode but the predictable outcome of the training regime itself.
Show all 6 sources
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- How AI Impacts Skill Formation
- Verification-Conditioned Use: A Qualitative Study on How Generative AI Reshapes Learning, Autonomy, and Market Entry for Junior Software Developers
- How much does AI impact development speed? An enterprise-based randomized controlled trial
- Exploring Student-AI Interactions in Vibe Coding
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- Anthropic Education Report: The AI Fluency Index
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot