People who feel more skilled after using AI often aren't — can we measure the gap between that feeling and real performance?
Can we measure perceived skill change against actual independent task performance?
This explores whether we can check how much people think their skills have changed after using AI against how well they actually perform once the AI is taken away, and what the corpus says about the gap between those two measures.
This explores whether self-reported skill change ("I've gotten better at this" or "I haven't lost anything") can be checked against what people can actually do without AI help. The short answer from the corpus is that the two are badly disconnected, and most current measurement setups can't see the side that matters. A pooled analysis of three studies found a correlation of only .055 between how competent people say they are with AI and how they actually perform. The confidence interval includes zero, which means self-ratings tell you essentially nothing about real ability Can self-ratings replace objective performance scores for AI competence?. So asking people whether AI has changed their skills doesn't work as a measurement method.
The obvious alternative is to watch what people actually do, but that has its own blind spot. Deployment telemetry, meaning the usage logs from AI tools in the workplace, records the work people produce with assistance. It never records what they could do alone. The corpus calls this a stock-formation gap: these systems see expertise being used, not expertise being built or lost. As a result, whether AI erodes or strengthens independent skill is still undetermined Can we measure whether AI erodes independent skill?. Answering the question takes deliberate unassisted tests, and routine data collection doesn't include them.
The less obvious finding is that the same perception-versus-reality gap shows up when humans judge AI models. Models trained to imitate ChatGPT convinced human evaluators that they had improved, because they copied its confident, fluent style. On new tasks, though, they were no more factual and generalized no better Can imitating ChatGPT fool evaluators into thinking models improved?. A skill-by-skill breakdown points the same way: surface skills like style level off early and are easy to copy, while reasoning keeps improving with scale and is much harder to fake Do all AI skills improve equally as models scale?. Instruction tuning shows a related pattern. Models trained on meaningless or even wrong instructions score about as well as models trained on correct ones, so what looks like task understanding is often just learning what the answer should look like Does instruction tuning teach task understanding or output format?.
Put together, this suggests a general rule. Fluent, polished output is what people perceive, and it can drift far from real capability whether the one being judged is a person or a model. That likely applies to people working with AI too: assisted work that looks good can make someone feel more skilled than they are. The fix looks the same in every case. Test on new tasks, without help, against an outside standard, the way benchmark studies test whether gains carry over to held-out problems Do AIDE2's improvements transfer to unseen tasks?.
One limit: the corpus shows that the gap exists and explains why current tools miss it. It doesn't yet include a study that tracks people's perceived skill change and their unassisted performance side by side over time. That experiment appears to be the missing piece.
Sources 6 notes
A pooled analysis of three studies found a correlation of only .055 between self-reported and objective measures of AI competence, with confidence intervals including zero. This provides no basis for substituting self-assessment for demonstrated performance.
Usage data registers assisted output but not independent capability. A stock-formation gap means current systems observe expertise in use better than expertise being built, leaving AI's skill effects fundamentally undetermined.
Imitation models fool human evaluators by mimicking ChatGPT's confident, fluent style while failing to improve factuality or generalization on novel tasks. The ceiling is set by base model capability, not fine-tuning method—better fundamentals, not shortcuts, drive real improvement.
FLASK's 12-skill decomposition reveals metacognition saturates at 7B parameters while logical efficiency plateaus at 30B, but reasoning and knowledge skills improve continuously. Open-source models successfully imitate surface-level style but fail at reasoning—confirming that distillation copies form not substance.
Models trained on semantically empty or deliberately incorrect instructions achieve comparable performance to those trained on full correct instructions, achieving 43% vs random baseline 42.6%. The semantic content of instructions appears largely irrelevant; what transfers is knowledge of the output space.
Show all 6 sources
The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap
- Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
- Available but Unclaimed: An Empirical Study of Human-AI Synergy
- UX Roundup (28 Sep 2026): Bogus Deskilling Research
- How AI Impacts Skill Formation
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
- The False Promise of Imitating Proprietary LLMs
- Do Models Really Learn to Follow Instructions? An Empirical Study of Instruction Tuning