SYNTHESIS NOTE
Topics›Knowledge After the Web›this note

Does removing AI tools actually measure real skill loss?

Nielsen questions whether lab experiments that take away AI and measure performance decline tell us anything useful about how AI affects workers in real jobs where the tool never gets removed.

Synthesis note · 2026-10-09 · sourced from Knowledge After the Web

Nielsen's complaint targets a specific experiment and the genre it belongs to. Shang Wu and colleagues (University of California, Irvine; HCOMP 2026) had 124 participants solve logic puzzles across three phases — no AI, optional AI, no AI again — and found that "everybody improved with practice, but AI users improved less, and their assisted scores overestimated their later unassisted scores." Nielsen calls this design "the confiscation study: hand people a tool ... let them use it, take it away, measure the decline, and conclude that the tool erodes skill." He grants "the measurement is valid," but argues "it rests on a false premise: outside the research lab, nobody confiscates the tool."

His reasoning runs by analogy: "We don't evaluate carpenters by confiscating their hammers and saws at the end of their apprenticeships and timing how long they take to drive a nail with a rock... The carpenter keeps the hammer. The knowledge worker keeps the AI." Because removal never happens in practice, he argues the field has been measuring the wrong dependent variable, and calls for a different program: "study what happens to users' higher-level skills after AI takes over the lower-level ones. Assume everybody has AI. Never measure performance without it... And distinguish among the ways people use AI, because their effects on learning are far from equal." He credits the Wu study with one finding that survives the design's flaw — skill gains tracked time spent reasoning before asking for help, not request frequency — but insists "no Bayesian model, however elegant, rescues a study that measures the wrong outcome."

This is a methodological argument, not a new measurement, and it sharpens Can we measure whether AI erodes independent skill? from the opposite direction: that paper's "stock-formation gap" says deployment telemetry can't see skill formation because it only watches assisted use with the tool left in place; Nielsen's complaint is that the lab-side alternative, removing the tool to test retention, watches a condition nobody actually lives in. Between the two, almost no clean window remains onto skill formation under permanent AI use. Can metacognitive feedback stop students from offloading to AI? is closer to the study Nielsen's program calls for: the assistant stays available throughout, and the intervention targets how answers get requested rather than whether access exists; its finding that heavy answer-requesters scored worse lines up with Nielsen's own point that help-seeking frequency, not help-seeking access, is what should be tracked. Do junior developers choose AI based on their ability to verify results? supplies the kind of distinction among usage modes that Nielsen says confiscation studies collapse into a single "used AI" variable.

The excerpt is commentary, not a study: Nielsen names the design flaw and proposes a direction but runs no experiment of his own, cites no body of confiscation-study results that reverse once the tool stays available, and does not specify which higher-level skills to measure or how. The implication he draws — that most published deskilling findings describe a condition absent from deployed use — is only as strong as the premise that the confiscation design can't be patched, a premise he asserts rather than demonstrates. What the excerpt does establish plainly is a standard worth applying to any deskilling claim encountered afterward: ask whether the tool was removed to produce the result.

Inquiring lines that read this note 5

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do AI coding tools measurably improve developer productivity and code quality? Does AI assistance help or harm professional skill development? How can AI systems reliably guide voters without introducing political bias? Does AI-assisted work increase total productivity or just shift time?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 102 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Nielsen argues the confiscation study measures the wrong outcome because nobody confiscates AI outside the lab