If you lean on AI for answers while practicing, do you do worse when it's taken away?
Does heavy reliance on AI answers during practice predict worse unaided test scores?
This explores whether students who lean on an AI for answers while practicing tend to do worse once the AI is taken away, and what the collection says about why.
This is about whether leaning on AI answers during practice costs you when you sit the test alone. The collection has one experiment that comes close. In a 704-person preregistered study, feedback that pointed out the cost of offloading cut answer requests to an LLM by half and raised unaided test scores by 51% Can metacognitive feedback stop students from offloading to AI?. Less reliance went with better unaided performance. That is an intervention result, not a measurement of who relies most and who scores worst. So "predicts" is a reasonable reading of the evidence but not something the study reports directly.
The same study has a twist. An effort-based reward, which pays people for trying harder, had no measurable effect on either offloading or test scores Can metacognitive feedback stop students from offloading to AI?. Motivation doesn't seem to be the problem. Students who offload aren't necessarily lazy. What worked was making the cost visible, which suggests they weren't seeing it on their own.
Other notes in the collection explain why the cost is hard to see. Fluent AI output makes people feel capable. Users read the smoothness of an answer as a sign of their own understanding, even though they didn't produce it Does processing ease mislead users about their own competence?. A broader account names four mechanisms that reinforce each other: it is unclear who did the work, answers feel fluent, thinking gets outsourced, and the pipeline is opaque. Together they make AI-assisted work look like your own skill How do AI tools trick users into overestimating their own skills?. Practicing with an AI can feel like learning while little is being retained.
This also means you can't ask students whether they're learning. In three pooled studies, self-rated AI competence correlated with objective performance at only .055, with a confidence interval that includes zero Can self-ratings replace objective performance scores for AI competence?. The unaided test may be the only honest signal, and it arrives after the habit has formed.
There is a parallel in how models are trained. Supervised fine-tuning can raise final-answer accuracy while step-by-step reasoning quality drops by 38.9%, so the model gets right answers without the inference that should produce them Does supervised fine-tuning improve reasoning or just answers?. A student copying correct answers from an AI during practice fits the same pattern, since the right answers hide the missing process. What the collection lacks is a study that tracks everyday learners across different amounts of AI use. For now the evidence is one strong experiment plus a well-supported account of why the harm goes unnoticed.