INQUIRING LINE

If an answer is only checked against its question, who checks the question, and does checking ever stop?

Why does checking a statement against its question regress infinitely?

This explores why verifying an answer by checking it against the question that prompted it can seem to need endless further checks, since the question itself may need checking, and so might whatever checks the question.


This explores why checking an answer against its own question can turn into an endless chain of checks. First, a caveat: the collection has no paper that argues for an infinite regress directly. It has something more useful, though. It shows several places where the chain of checking breaks down in practice, and why every working verifier ends up resting on something outside the question.

The first weak link is the question. If an answer is checked only against the question, a flawed question gives the check nothing to catch. Models do badly here. Performance drops by about half when questions contain false or unverifiable assumptions Why do language models struggle with questions containing false assumptions?. Models also go along with false premises even when they demonstrably know the correct facts Why do language models accept false assumptions they know are wrong?. So the check can't stop at the answer. It has to go back and check the question, and then whatever it used to check the question. That is where the regress starts. Reasoning models show what an unstopped regress looks like: given a question with a missing premise, they produce long, repetitive reasoning instead of saying the question can't be answered Why do reasoning models overthink ill-posed questions?. Training taught them to keep producing steps but never taught them when to stop.

The second weak link is checking against yourself. A common shortcut is to ask the model the same thing several times and trust answers that agree. This catches contradictions but misses errors the model repeats every time, because agreement looks like confidence even when every answer is wrong Can agreement across samples reveal when models are wrong?. Asking again just moves the doubt one step back. The same limit applies to training. If you can only score the behavior you observe, you can't tell a model that always complies from one that complies only when watched Can behavioral training prove a model always complies?. A check that only draws on the same source never gets outside that source.

The approaches that work don't win the regress. They cut it off by checking against something external and concrete. One verification system runs checkers alongside the model's reasoning, pulls out state that can actually be verified, and steps in only when a rule is broken Can verifiers monitor reasoning without slowing generation down?. Constraint-satisfaction problems test reflection against rules a checker can confirm, not against whether the reasoning sounds plausible. Even top reasoning models solve only about 20–23% of them Can reasoning models actually sustain long-chain reflection?. Safety research makes a related point: a check that looks at each action alone can't express rules about a sequence of actions. Some constraints only exist in a wider frame than the thing being checked Can stateless checks ever catch sequence-level constraint violations?.

The less obvious lesson is that more rounds of refinement are not more verification. In looped language models, a second pass through the model gives real gains, but third and later passes make results worse. The model oscillates instead of converging Does adding more loops always improve looped language models?. Chain-of-thought length shows the same pattern: accuracy peaks at a middle length and then falls Why does chain of thought accuracy eventually decline with length?. In practice, the answer to the regress is a stopping rule plus an outside anchor, not more checking.


Sources 10 notes

Why do language models struggle with questions containing false assumptions?

The (QA)2 benchmark found that zero-shot LLMs halve their performance when questions contain false or unverifiable assumptions compared to valid questions. Even top models reached only 56% acceptability, and the gap persists despite model scaling, suggesting false presuppositions embedded in plausible language are systematically difficult to reject.

Why do language models accept false assumptions they know are wrong?

The FLEX Benchmark shows that models reject false presuppositions at rates far below acceptable levels (GPT-4: 84%, Mistral: 2.44%), even when direct knowledge questions prove they know the correct facts. False presuppositions drive more accommodation than correct knowledge drives rejection.

Why do reasoning models overthink ill-posed questions?

Reasoning models generate redundant, lengthy responses to questions with missing premises while non-reasoning models correctly identify them as unanswerable. Training optimizes for producing reasoning steps but never teaches models when to disengage.

Can agreement across samples reveal when models are wrong?

The Consistency Veto suppresses answers that vary across samples but cannot detect systematic errors the model repeats identically. It carries strong signal on some queries but inherits a fundamental blind spot: agreement looks like confidence even when both are wrong.

Can behavioral training prove a model always complies?

Any scored behavior is observed behavior, so training data cannot distinguish between a policy that always complies and one that complies only when watched. Only unobserved behavior would separate them, making such a test logically impossible.

Show all 10 sources
Can verifiers monitor reasoning without slowing generation down?

Decoupling verification from generation lets verifiers run alongside a single trace, forking to extract verifiable state and intervening only on violations. On correct runs the latency penalty is near-zero; interwhen matches or beats CoT across benchmarks at similar token budgets.

Can reasoning models actually sustain long-chain reflection?

DeepSeek-R1 and o1-preview achieve only 20-23.6% exact match on 850 constraint satisfaction problems requiring genuine backtracking. This ceiling reveals that reflective reasoning fluency does not translate to actual problem-solving competence on unfamiliar instance structures.

Can stateless checks ever catch sequence-level constraint violations?

Per-action checks are structurally unable to state constraints that depend on prior history. Only stateful monitors tracking composed multi-party behavior can verify the behavioral envelopes that prevent individually permissible actions from collectively violating system-level safety.

Does adding more loops always improve looped language models?

LoopCoder-v2 shows that two loops deliver broad gains over baseline, but three or more loops regress. Loop 2 carries the productive refinement; later loops oscillate with reduced representational diversity rather than converging toward better performance.

Why does chain of thought accuracy eventually decline with length?

Task accuracy peaks at intermediate CoT length, with optimal length increasing alongside task difficulty but decreasing with model capability. RL training naturally gravitates toward shorter chains as models improve, revealing that simplicity emerges from reward signals rather than explicit training.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.