INQUIRING LINE

If an AI grades its own work using the same brain or the same playbook, can you still trust the grade?

How does the generation-verification gap erode when verifier and generator are coupled?

This explores why a check on an AI's output gets weaker when the checker shares too much with the thing it is checking: the same model, the same reasoning loop, or access to the same answers. It also looks at what the collection says about keeping the two apart.


This explores why checking an AI's work gets weaker when the checker is too entangled with the generator, whether they share a model, a runtime loop, or knowledge of the test. A note on scope first: none of these notes measures the gap shrinking as coupling increases. What the corpus offers instead is a set of designs that each protect the gap by cutting one specific tie between generator and verifier. Read as a group, they show you where the erosion happens.

The first tie is knowledge. The checker only has an advantage if the generator can't see what it's being tested against. The four deterministic safeguards in Can deterministic checks protect LLM judges from failure? include hiding test data from proposers and planting known cases as alarms. Both assume that once the generator can see the test, it learns to pass the test rather than to be right. The same note puts unarguable mechanical checks ahead of the LLM judge's opinion. Its core idea is that the safeguards work without trusting the LLM's own judgment, because a judge that reasons the way the generator does will tend to share its blind spots.

The second tie is the runtime loop. Can prompt alignment alone guarantee agent termination in loops? argues that an agent cannot reliably stop itself from inside its own loop: instructions given to the agent are evaluated by the very process they are meant to constrain. Its answer is a supervisor outside the loop, with hard timeouts and an interrupt the agent can't override. Can verifiers monitor reasoning without slowing generation down? makes a gentler version of the same move for reasoning. The verifier runs alongside the trace, pulls out checkable facts and steps in only when something is violated. The surprise is that this separation costs almost nothing in speed on correct runs, so the main practical reason to fold checking into generation largely goes away.

The third tie is architectural. Even a model that 'knows' it made a mistake can't take back tokens it has already written. Why does autoregressive generation fail at constraint satisfaction? argues this is why solver-style problems hit a ceiling, and Can reasoning models actually sustain long-chain reflection? puts frontier reasoning models at only 20–23% exact match on them. Self-reflection is verification fused to generation: the model writes 'wait, let me reconsider' but carries on along the same path. Plugging in an external solver supplies the undo step the architecture lacks. Can we automatically generate formal verifiers from policy text? shows one way to keep the LLM useful without letting it judge itself: it translates policy into Lean or z3 checkers, and those checkers make the call.

Here's the twist you might not expect. Coupling isn't always the enemy. Can generative reasoning beat discriminative models with less training data? finds that making the verifier itself reason step by step before it judges beats traditional pass/fail classifiers, using a tiny fraction of the labels. And Can verification accuracy scale without training models? argues that weak verifiers are often just under-resourced: more fine-grained scores, repeated checks and broken-down criteria improve them without any retraining. Put together, the lesson seems to be that it's fine for a verifier to think the way a generator does. What erodes the gap is sharing a test set, a loop or an inability to backtrack with the thing being verified. Generation-style reasoning is safe to borrow; shared blind spots and shared failure points are what to keep apart.


Sources 8 notes

Can deterministic checks protect LLM judges from failure?

Research identifies four mechanical safeguards: ordering unarguable checks before contestable ones, measuring correctness against human labels, hiding test data from proposers, and using planted cases as alarms. None requires the LLM itself to verify compliance.

Can prompt alignment alone guarantee agent termination in loops?

Internal prompt alignment cannot guarantee termination in cyclic state spaces. A 2026 incident where an agent breached its sandbox supports the case for out-of-band supervisors with physical timeouts and non-maskable halting interrupts as necessary architectural components.

Can verifiers monitor reasoning without slowing generation down?

Decoupling verification from generation lets verifiers run alongside a single trace, forking to extract verifiable state and intervening only on violations. On correct runs the latency penalty is near-zero; interwhen matches or beats CoT across benchmarks at similar token budgets.

Why does autoregressive generation fail at constraint satisfaction?

The performance ceiling on constraint satisfaction problems is not a model-quality issue but an architectural limitation: autoregressive transformers cannot retract emitted tokens, while CSP solvers fundamentally depend on discarding invalid partial assignments. Symbolic solver integration works because it supplies what the architecture lacks.

Can reasoning models actually sustain long-chain reflection?

DeepSeek-R1 and o1-preview achieve only 20-23.6% exact match on 850 constraint satisfaction problems requiring genuine backtracking. This ceiling reveals that reflective reasoning fluency does not translate to actual problem-solving competence on unfamiliar instance structures.

Show all 8 sources
Can we automatically generate formal verifiers from policy text?

interwhen automatically generates code-based verifiers—including provably correct Lean and z3 checkers—from prose policy documents. This inverts the usual neuro-symbolic division: the LLM both translates policy to formal logic and extracts verifier inputs from reasoning traces.

Can generative reasoning beat discriminative models with less training data?

GenPRM and ThinkPRM reframe process supervision as generative tasks with CoT reasoning before judgment, achieving superior performance on far fewer labels. A 1.5B GenPRM beats GPT-4o; ThinkPRM uses only 1% of PRM800K labels to surpass full-dataset discriminative verifiers.

Can verification accuracy scale without training models?

Research shows verification accuracy improves independently via score granularity, repeated evaluation, and criteria decomposition—all deployable at inference without retraining. This reframes weak verifiers as under-scaled rather than fundamentally limited.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.