If AI quietly erodes your skills, how long would you have to watch to notice, and would a calendar even help?
What timescale matters for measuring gradual skill loss in augmented work?
This explores what time window you would have to watch to catch people gradually losing skills while working with AI, and whether calendar time is even the right clock.
This explores what time window you would have to watch to catch people gradually losing skills while working with AI, and whether calendar time is even the right clock. The corpus has no long-term study of human skill decay, so it can't say 'six months' or 'two years'. It does suggest that the right clock is how often you check the person without the AI, and that the metric you pick can hide a slow decline.
The key finding is that AI-enhanced abilities work like an exoskeleton. Workers produce skilled-looking output while the AI is present and revert to baseline when access is removed (Does AI assistance build lasting skills or temporary abilities?). So watching performance during augmented work can't show skill loss, however long you watch, because the output looks the same either way. Measurement has to include unaided checkpoints, and the gaps between them set your resolution. The note also complicates the question. If the ability never persisted independently, there may be nothing gradually leaking away, just an unchanged baseline hidden behind boosted output. Whether that baseline itself sinks over months of heavy reliance is the part this corpus doesn't test.
Whether a decline looks gradual or sudden also depends on how you score it. Work on LLM 'emergent abilities' found that sharp, unpredictable jumps vanish when you swap an all-or-nothing metric for a continuous one, and the same outputs show smooth improvement (Are LLM emergent abilities real or measurement artifacts?). That result is about models, not people, but the lesson transfers by analogy. If you score a person's skill as pass/fail on a task, slow erosion looks like nothing for a long stretch and then a cliff. A graded measure would show the slope much earlier, so a finer metric shortens the timescale you need.
The workflow itself may be a better unit than the calendar. In DELEGATE-52, frontier models degrade documents through subtle corruption that keeps the surface looking intact, while weaker models visibly delete content, which makes frontier failures harder to detect at workflow scale (Does model capability change how documents degrade?). Errors that accumulate silently across many delegated steps only show up over many handoffs. It also means the human checking skill you'd want to measure matters most when errors are hardest to see. In that setting the useful clock is how many delegation cycles have passed since anyone verified the work unaided.
So the answer, as far as the corpus goes, is to measure in withdrawals and delegation cycles rather than weeks, and to use a graded metric so slow change shows up before it becomes a cliff. The gap is real: nothing here follows the same people over months to say how fast an unaided baseline erodes, if it does.
Sources 3 notes
Research shows AI assistance creates temporary capability extensions—workers produce skilled-looking output while AI is present but revert to baseline performance when access is removed. This differs fundamentally from true skill, which persists independently.
Sharp, unpredictable capability transitions vanish when using continuous metrics instead of discontinuous ones. The same model outputs show smooth predictable improvement with scale, suggesting emergence is a measurement choice rather than a real behavioral change.
DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- LLMs Corrupt Your Documents When You Delegate
- Are Emergent Abilities of Large Language Models a Mirage?
- Are Emergent Abilities in Large Language Models just In-Context Learning?
- Progress Measures For Grokking Via Mechanistic Interpretability
- AI Assistance Reduces Persistence and Hurts Independent Performance
- Large Language Model Reasoning Failures
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- How AI Impacts Skill Formation