Does AI help or harm learning based on how it's designed?
Two randomized trials with overlapping research teams tested whether the same AI technology improved or worsened student math and programming outcomes. The difference turned on a single design choice: whether AI gave answers directly or tutored students through problems.
Ethan Mollick argues that whether AI helps or harms learning turns on a narrow design choice: does it give the answer, or does it tutor. He cites "two papers with an overlapping research team (including peers at Wharton)." In the first, roughly a thousand high school students in Turkey learned math either with plain ChatGPT or with no AI access; the ChatGPT group "did their homework better and thought they were learning more," but "at test time, they underperformed their classmates without ChatGPT." In the second, close to a thousand students across ten Taipei high schools took a five-month Python course; those given "a personalized sequence of problems by an AI tutor scored 0.15 standard deviations higher on a final exam taken without AI help," which Mollick says some estimates put at "six to nine months of additional schooling."
The mechanism Mollick gives is effort substitution: plain ChatGPT "was really just giving them answers, and actual learning requires mental effort. By short-circuiting effort, you short-circuit learning." The Taipei tutor instead sequenced problems to the student without removing the work itself. He extends the same mechanism outside the classroom to a BCG field experiment he co-ran with Fabrizio Dell'Acqua and others on 758 consultants: those with GPT-4 access "vastly outperformed" on most tasks, but on one task built to defeat the model, "the same elite consultants who outperformed on everything else" mostly failed to catch its confident wrong answer. He names the underlying failure mode "cognitive surrender," a term he credits to Wharton colleagues, and argues the Taipei result points toward system-level design constraints rather than user willpower as the fix — a lever he says "we don't see much of ... in the consumer products."
This sharpens Does AI assistance actually harm the way developers learn? by supplying two further randomized trials, with effect sizes, for the same split between delegation-like use (answers) and engagement-preserving use (tutoring). It also bears on Can metacognitive feedback stop students from offloading to AI?: both that study and Mollick's reading of Taipei treat the fix as something built into the system's design rather than left to the user's self-discipline, though Mollick frames this as a bet about commercial AI products in general rather than a tested feature.
The piece is a synthesis of other teams' studies, not primary data Mollick collected himself, and its numbers are given loosely — "about a thousand," "close to a thousand," and a "six to nine months" conversion attributed to unnamed "estimates" rather than derived in the piece. The Turkey and Taipei results also come from two specific populations (one math, one programming) and age range, so the answer-versus-tutor distinction is a design principle worth testing elsewhere, not yet a general law of AI-assisted learning. The implication Mollick draws, that this is "going to be a defining challenge of the coming years" solved mostly by product design choices rather than user intentions, follows only as strongly as that evidence allows.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does AI assistance help or harm professional skill development?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does AI assistance actually harm the way developers learn?
When developers use AI tools while learning new programming concepts, does it impair their ability to understand code, debug problems, and build lasting skills? Understanding this matters for how we deploy AI in education and training.
same answer-vs-engagement divide, here with two further RCTs and effect sizes
-
Can metacognitive feedback stop students from offloading to AI?
When learners practice with an AI assistant, does making them aware of the downsides of offloading their work reduce how much they ask the AI to solve for them? And does that change improve their performance on tests without help?
both treat system-level design, not user willpower, as the lever against offloading
-
Does ChatGPT harm informal learning compared to Google Search?
An 8-day experiment tested whether using ChatGPT for self-directed learning produces different knowledge gains than Google Search, and explored what mechanisms might explain any differences.
B extends A's Turkey finding: ChatGPT-as-answer-tool again harms learning, via diminished agency and information distortions
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Choosing to Stay Human
- AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting
- Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap
- UX Roundup (28 Sep 2026): Bogus Deskilling Research
- Available but Unclaimed: An Empirical Study of Human-AI Synergy
- AI Assistance Reduces Persistence and Hurts Independent Performance
- AI Meets the Classroom: When Does ChatGPT Harm Learning?
- AI-written admissions essays are widespread but penalized
Original note title
Mollick contrasts a Turkey ChatGPT math trial that lowered test scores with a Taipei AI-tutor trial that raised them 0.15 SD