SYNTHESIS NOTE
Topics›Domain Specialization›this note

Can we measure whether AI erodes independent skill?

Current telemetry tracks how people use AI but not whether they become more capable without it. Existing measurement tools cannot yet determine if AI helps or hurts skill formation at scale.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

Toward Measuring AI's Effects on Skill Formation argues that the instruments watching AI use are lopsided. Deployment telemetry observes "tasks, interaction patterns, and outputs" but not "whether users become more capable of performing those tasks independently." The Anthropic Economic Index, built from over four million assistant conversations, shows AI use concentrated among skilled professionals, and the paper's point is what that kind of data cannot see. Controlled learning experiments measure independent capability more directly, but in narrower populations. The authors call the mismatch a stock–formation measurement gap: current systems "observe the use of existing expertise more readily than the formation of future expertise." The claim is deliberately limited: "The claim is not that AI has been shown to erode skill formation at population scale. It is that existing measurement cannot determine whether it does."

The mechanism rests on four of the paper's seven quantities. Z is the assigned condition, the interface or defaults, running from answer delivery through hints to evaluation of one's own attempt. A is the realized allocation of cognitive work, a profile over functions such as planning, generation, monitoring and verification. Y is the immediate output, and ∆K is the change in unassisted capability, measured with the tool removed. Z shapes A without determining it, since a hints interface can be used passively. Because the effortful parts of a task "are part of the mechanism of formation," an assistant that absorbs them by default absorbs formation along with the friction, so effects on ∆K "cannot be read off improvements in Y." The paper's Tier 1 covers three randomized trials and reads Z as moving ∆K in both directions. In one of them, a base GPT-4 tutor raised assisted practice by 48% while lowering unassisted exam scores by 17%, a result Bastani et al. (2025) measured and the paper reports.

The nearest notes are instances of this gap seen from different sides. Does AI assistance help workers learn lasting skills? reports a measured case in which a within-study gain did not carry into later independent work, which is the Y-versus-∆K split in concrete form; this excerpt adds why deployment telemetry would miss that outcome. Can metacognitive feedback stop students from offloading to AI? is an intervention on the handover, the A-side lever, and the paper's program would evaluate it through unaided retention rather than test-day scores alone. The institutional version appears in Do university AI policies actually protect what credentials mean?: boundaries are stated more clearly than the evidence standards that would show what a credential certifies, which is the same measurement failure applied to assessment.

The excerpt establishes less than its framing suggests. It calls itself "a motivated narrative synthesis, not a systematic review," performs no risk-of-bias grading, and notes that the assigned condition was randomized in only three studies and the realized allocation in none. The deployment data is "a descriptive illustration of the gap," and the research program it proposes, which links consented usage records to independent assessments while varying answers, hints, feedback and evaluation, is a design rather than a result. What follows at the strength the evidence allows is narrow: whether sustained AI use changes independent capability is a real, measurable question that deployed instruments cannot currently answer. It does not follow that AI is eroding learning, and the paper itself warns that unfounded alarm carries costs of its own.

Inquiring lines that read this note 10

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does AI assistance help or harm professional skill development? Does AI-assisted work increase total productivity or just shift time? Does AI deployment reduce or exacerbate workplace inequality and income instability?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 91 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

deployment telemetry observes expertise in use but not expertise forming — the stock-formation gap leaves AI's skill effects undetermined