INQUIRING LINE

AI-assisted work can drown both the AI and the human in too much plausible-looking output — can careful, built-in checking stop that?

Can System 2 oversight prevent information overload in LLM-assisted work?

This explores whether slow, deliberate checking (the 'System 2' kind of thinking, done by a person or built into the workflow) can keep LLM-assisted work from burying people, or the model itself, in more material than anyone can properly check.


This explores whether slow, deliberate checking can keep AI-assisted work manageable when the volume of output outruns anyone's ability to review it. The corpus has no study of human cognitive overload as such. What it does show is surprising: the overload problem runs in two directions. The model gets overloaded by too much context, and the human gets overloaded by too much plausible-looking output. Deliberate oversight helps with both, but only when it is built into the structure of the work. Asking a reviewer to stay alert does not do it.

On the model's side, the most effective form of 'System 2' turns out to be ordinary software design. Can algorithms control LLM reasoning better than LLMs alone? describes wrapping LLM calls inside an explicit algorithm that controls the flow of the task. Each step sees only the information it needs, and the rest is deliberately hidden. Can reasoning and tool execution be truly decoupled? makes a related move: the model plans first and fetches tool results afterward, so the prompt doesn't swell with every intermediate observation. In both cases the deliberate layer prevents overload by withholding information. Adding more review on top would not have the same effect.

On the human side, the news is sobering. Do frontier LLMs silently corrupt documents in long workflows? found that even the strongest models degrade about a quarter of a document's content across long chains of handoffs, and that spot checks miss it. Does model capability change how documents degrade? explains why: weaker models visibly delete things, while frontier models quietly change things so the surface still looks intact. A careful reader catches missing paragraphs. A careful reader rarely catches subtly wrong ones. So the better the model, the less a human's deliberate attention actually buys you. You also can't hand the checking back to the model. Do large language models fabricate user attributes beyond available evidence? found that models which rate themselves as more careful are actually less careful, and What limits autonomous capability in large language models? describes a formal limit on self-checking: a model can only improve as far as it can reliably tell its good outputs from its bad ones.

What works in practice is to narrow where human judgment gets applied. In Can LLMs generate workflows without touching proprietary data?, the model assembles workflows from vetted building blocks, so a person reviews a short, readable plan instead of raw output. How can conferences detect and handle LLM misuse in peer review? shows the same idea at conference scale. Automated flags went to human reviewers as one input among several, and hard enforcement was reserved for the one thing that is easy to verify: whether the cited references actually exist. Timing matters too. Why do AI assistants get worse at longer conversations? shows that models lock in early guesses, so oversight at the start (clarifying the request before the model commits) is worth more than review at the end.

The short answer is yes, but not as more careful reading. Deliberate oversight prevents overload when it is designed into the workflow: hiding irrelevant context from the model, limiting it to vetted actions, and placing human checkpoints where errors are cheap to verify. The finding you may not have expected is that model quality makes the problem harder. As models improve, their mistakes get quieter, and careful human attention becomes less effective as a safety net.


Sources 9 notes

Can algorithms control LLM reasoning better than LLMs alone?

LLM Programs embed LLMs within explicit algorithms that manage control flow and state, presenting only step-specific context to each LLM call. This information hiding addresses capability and context window limits while treating complex reasoning as modular, debuggable sub-tasks.

Can reasoning and tool execution be truly decoupled?

ReWOO and Chain-of-Abstraction both decouple reasoning from tool responses through different mechanisms—planning-before-execution and abstract placeholders respectively—eliminating quadratic prompt growth and sequential latency while maintaining reasoning quality.

Do frontier LLMs silently corrupt documents in long workflows?

Even the strongest models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) degrade documents by ~25% over long relay workflows across 52 domains. Degradation decelerates but never plateaus, and errors compound silently, remaining undetected in spot-checked outputs.

Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

Do large language models fabricate user attributes beyond available evidence?

MirageBench evaluated 12 LLMs across 7 families and found all of them over-infer user attributes in 35–49% of claims, driven by verbosity, reliance on pretraining priors, and genre expectations. Models that self-assess as over-inferring less actually over-infer more when judged independently.

Show all 9 sources
What limits autonomous capability in large language models?

Multi-agent deliberation produces specific failure modes (Degeneration-of-Thought, Silent Agreement), alignment at scale includes problematic self-valuation, and self-improvement is formally bounded by the generation-verification gap. Measurement error and conditional compliance hide the true capability ceiling.

Can LLMs generate workflows without touching proprietary data?

FlowMind demonstrates that LLMs can generate on-the-fly workflows for spontaneous tasks by orchestrating calls to vetted APIs rather than accessing data directly, eliminating confidentiality risks while maintaining high-level human inspection and feedback.

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

Why do AI assistants get worse at longer conversations?

LLMs perform at 90% accuracy with single-message instructions but drop to 65% across natural conversation. Models lock into early guesses when information arrives gradually and cannot course-correct, a behavior induced by RLHF training that rewards helpfulness over clarification.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.