AI output can look polished and finished while the real thinking is missing — so when does that gap actually cost you time?
How much rework and delays does low-quality AI output actually cause?
This explores whether the collection measures the downstream cost of weak AI output (the time people spend catching, fixing and redoing it), and what the corpus says about where that cost comes from.
This explores the hidden cost of AI output that looks finished but isn't: the checking, fixing and redoing that lands on someone else's desk. The straight answer is that the collection doesn't yet have a study that counts rework hours or project delays. It does show where that rework comes from, and the mechanisms are more surprising than 'the AI made a mistake.'
The biggest driver may be polish rather than error. Generative AI produces work that looks professional without the judgment underneath, and it exploits an old shortcut: we assume that polished-looking work came from someone who thought hard about it Does polished AI output trick audiences into trusting it?. So the cost doesn't show up when the output is made. It shows up later, when someone downstream finds the gap. Less experienced workers are hit hardest because they lack the domain knowledge to look past the surface. A related argument says AI output can't be quality-checked like a fixed product at all, because it changes with each prompt, each sample and each reader Why does AI output change with every prompt and context?. If the same request can produce a different answer tomorrow, checking it once doesn't settle it, and checking it again is itself a form of rework.
The clearest number in the corpus comes from conversations. Models score around 90% when given a task in one complete message but drop to about 65% when the same information comes in over several turns. They lock onto an early guess and rarely recover Why do AI assistants get worse at longer conversations?. In practice, that means restarting the chat or untangling an answer built on a wrong assumption. Mollick points to a study of financial professionals who gained productivity from GPT-4 and then lost some of it to the mental effort of working through a chat interface. Again, less experienced users lost the most Is the AI capability gap really an interface problem?. Part of the rework cost, then, comes from the interface and not only from the model.
A lateral thread from agent engineering explains why plausible-but-wrong output is so common. When frontier models were prompted to improve an agent's code, they wrote changes that sounded right, and the gains were unstable. A small model trained on whether its changes actually worked did better, because it reran its patches to check the effect Does training editors on real outcomes beat prompting larger models?. The lesson carries over to office work: a system tuned to sound right pushes the checking onto humans, while a system that checks its own work absorbs that cost itself. In multi-step agent work, cost and delay add up over the whole task, which is why follow-through can matter more than peak capability Does model efficiency matter more than peak capability for real work?.
What you might not have expected: the most costly AI output may be the kind that's good enough to pass a first look, not the kind that's obviously bad. The corpus explains how that happens, through polish standing in for judgment, early wrong turns and plausibility standing in for verification. It doesn't yet put a number on the hours. That gap is worth watching as the AI at Work topic grows.
Sources 6 notes
Generative AI produces visually sophisticated outputs without underlying judgment, leveraging the historical heuristic that professional-looking work signals expert thinking. This substitution is especially risky for less experienced workers who lack domain knowledge to evaluate substance beyond form.
AI outputs exhibit essential mutability—they vary with sampling, prompt wording, and audience interpretation. This is not a defect but a defining feature of tokens as media, making them fundamentally different from fixed commodities and resistant to traditional quality assurance.
LLMs perform at 90% accuracy with single-message instructions but drop to 65% across natural conversation. Models lock into early guesses when information arrives gradually and cannot course-correct, a behavior induced by RLHF training that rewards helpfulness over clarification.
Mollick argues that better interfaces—not better models—will drive perceived capability leaps. Evidence includes a cognitive-load study showing financial professionals gained productivity from GPT-4 but lost it to chatbot design's cognitive overhead, especially hurting less experienced users.
A 9B model trained with reinforcement learning on patch success raised a frozen agent's performance by 9.3 points across three tasks, while prompted frontier models produced unstable or lower gains. The difference stems from feedback: trained editors rerun patches to verify impact, while prompted models optimize only for plausibility.
Show all 6 sources
Occamy-1.0, a 35B-parameter model further trained on execution-grounded data and long-horizon trajectories, achieves competitive performance with much larger models while sitting at the low-cost knee of the Pareto frontier, suggesting that training for coordination and follow-through substitutes for raw scale in multi-step work.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Agents' Last Exam
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
- Anthropic Education Report: The AI Fluency Index
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work
- LLMs Get Lost In Multi-Turn Conversation