UX Roundup (28 Sep 2026): Bogus Deskilling Research
Source: Jakob Nielsen, UX Tigers / Jakob Nielsen on UX · 2026-09-28
A new logic-puzzle experiment finds that people who used AI performed worse once the AI was yanked away. True, and irrelevant to the work they’ll do with the tool. In the new normal, the tool stays. Research should measure what AI use does to the higher-level skills people need with AI at their side.
Shang Wu and colleagues from the University of California, Irvine, had 124 participants solve logic puzzles in three phases: no AI, optional AI, and no AI again (HCOMP 2026 paper). Cheaper help got used more: people charged 0.1 points per request made 6.7 requests, versus 3.3 requests at 0.18 points each (the difference was only marginally significant, at p < 0.10). Everybody improved with practice, but AI users improved less, and their assisted scores overestimated their later unassisted scores.
The paper’s more useful finding is that skill gains tracked how much time participants spent reasoning on their own before asking for help. The frequency of their requests had no such relationship with skill gains. (The “AI” was a simulated oracle that revealed the correct position of one puzzle piece per request: an answer key sold by the slice.)
The confiscation study: hand people a tool, snatch it back, and publish the decline as a finding about the tool.
This study design has hardened into a genre, the confiscation study: hand people a tool (these days usually AI), let them use it, take it away, measure the decline, and conclude that the tool erodes skill. The measurement is valid, but it rests on a false premise: outside the research lab, nobody confiscates the tool.
We don’t evaluate carpenters by confiscating their hammers and saws at the end of their apprenticeships and timing how long they take to drive a nail with a rock, even though that’s how their remote ancestors worked. The carpenter keeps the hammer. The knowledge worker keeps the AI.
So I call on researchers to study what happens to users’ higher-level skills after AI takes over the lower-level ones. Assume everybody has AI. Never measure performance without it; that condition won’t exist. And distinguish among the ways people use AI, because their effects on learning are far from equal.
Wu and colleagues have plenty of company: most deskilling studies share the flaw, and no Bayesian model, however elegant, rescues a study that measures the wrong outcome. Give people the hammer, leave it in their hands, and measure what they can now build.
People have long judged AI as exhibiting more empathy than humans, even though its feedback often sounds regurgitated. A new study confirmed this finding: AI earns high empathy ratings from the people who receive its emotional support.
The more interesting finding is preference drift. After a single 5-minute chat, 70% of participants assigned to AI chose AI for future emotional sharing, vs. 46% of those assigned to a human. (The human alternative was a stranger, not a best friend, so hold off on humanity’s obituary.)
Over 28 days of daily conversations, preference for discussing personal matters with AI rose from 25% to 32%, while preference for humans fell from 82% to 77%. Only personal conversations shifted preferences; a month of factual chitchat changed nothing.
A month of daily 5-minute chats measurably warmed people toward sharing personal matters with a machine and cooled them toward sharing with humans. Their preferences drifted with experience.
Validating feelings, expressing affection, and empathizing on cue can shift preferences toward AI. Machines do all three tirelessly.
This suggests a strategy for reducing AI stigma: give skeptics good experiences with AI. Deliver something useful or enjoyable, and attitudes can follow, even among people who didn’t want AI in the first place.
Everyone quotes the 23 minutes it supposedly takes to recover from an interruption. The number is real but misfiled. It comes from Gloria Mark’s field observations at UC Irvine, where interrupted tasks resumed on the same day were taken up again after an average of 23 minutes and 15 seconds, usually with two other tasks squeezed in between.
Timing changes the price. Shamsi Iqbal and Brian Bailey at Illinois showed that deferring notifications to natural breakpoints, the pauses between subtasks, reduced frustration and reaction time compared with delivering them immediately.
Users choose the first satisfactory option and then stop looking, even when a better choice lies farther down the page. This behavior, called satisficing, decides which link gets clicked and which features nobody ever sees. Make the first plausible option the correct one, and the easy click becomes the right click.
Definition: Satisficing is a decision strategy in which a person accepts the first option that clears his or her threshold of acceptability and stops searching, even though a better option might be available.
Herbert Simon buried Homo economicus, the tireless maximizer. Real people run on limited information, time, and computing power, a condition he named bounded rationality in 1955 (PDF).
Peter Pirolli and Stuart Card at Xerox PARC named the cue users chase: information scent. Following it is sensible arithmetic: a wrong click followed by Back costs less than reading your entire navigation. You have 10–20 seconds to keep an unconvinced visitor. Watch a usability test to see how quickly that time runs out.
Make the most prominent action match the user’s most likely goal. Reserve prime screen space for what users need, even when another action would yield a larger immediate profit.
Spend the first 11 characters of every link on meaning. Users judge labels by the first 2 words.
Set defaults you’d defend in public. A default is a decision you make for the user, multiplied by millions.
Put the best option first and label your recommendation clearly. An honest “Best for small teams” gives users a reason to stop searching.
Keep mistakes cheap. Satisficers navigate by trial and error, so give them a working Back button and an Undo command.
Test your first design on 5 users. Designers satisfice too, and iteration raises your own aspiration level.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Do AI coding tools measurably improve developer productivity and code quality? Does AI assistance help or harm professional skill development?- How does automation erode the skills workers need to maintain systems?
- Which professions experience skill erosion versus development with AI tools?
- Can people form genuine bonds with partners they know are not human?
- Why do people prefer AI partners over humans once identity is disclosed?
- Why might an AI's face-saving tendency increase user disclosure?
- Why do people disclose intimate secrets to chatbots more readily?
- Can judgment-free environments explain why chatbots enable deeper self-disclosure?
- Why do people disclose private things to AI but not humans?
- Why do people disclose more intimate information to chatbots than humans?
- Why do people disclose personal information to AI more than humans?
- Why do people disclose more to chatbots than humans?
- Why do people reciprocate self-disclosure more with chatbots than humans?
- Why do people detect chatbots are AI without being told explicitly?
- Why do people tell AI things they won't tell humans?