INQUIRING LINE

A simple pass/fail test can work surprisingly well — but only if the thing it's checking stays frozen and unchanging while you test it.

What makes the minimal-criterion check effective within a fixed representation?

This explores why a simple pass/fail check (a 'minimal criterion' that only asks whether something clears a bar, not how well it scores) works when the underlying representation is held fixed rather than being retrained alongside the check.


This explores why a bare pass/fail check works when the thing being checked sits on a fixed, frozen representation. The retrieved notes don't include any material on minimal-criterion methods themselves, such as minimal-criterion novelty search or coevolution, so this can't be answered directly. What the collection does have is adjacent work on three related questions: when a cheap check against a fixed substrate is trustworthy, when it misleads, and what makes it sharper.

The strongest case for checks on a frozen base comes from work that leaves the model untouched and acts only on its internal states. Representation finetuning edits the frozen hidden representations directly instead of updating weights. It beats LoRA while using 10–50x fewer parameters Can editing hidden representations beat weight updates for finetuning?. The lesson carries over to checks: when the representation is held fixed, a small, low-dimensional intervention or test can do a lot, because it isn't chasing a moving target. Verification research points the same way. Checks get more accurate through finer scoring, repeated evaluation and breaking criteria into parts, all at inference time with no retraining Can verification accuracy scale without training models?. A weak check is often just an under-scaled one.

The catch is that a fixed representation can pass a minimal bar while being badly organized inside. Models can contain every feature a linear probe needs, with perfect accuracy, while their internal structure is fractured and fragile under shifts in the data Can models be smart without organized internal structure?. A sibling result on reasoning shows the same thing: most models 'pass' constraint problems by defaulting to the harder option, and they drop up to 38.5 points when the constraints are removed Are models actually reasoning about constraints or just defaulting conservatively?. A threshold check confirms the bar was cleared. It doesn't tell you why it was cleared.

So what makes a check effective is what it looks at, not how strict it is. A small verifier that reads full token-to-token similarity maps reliably catches structural near-misses that a compressed single-vector comparison lets through Can verification separate structural near-misses from topical matches?. The same fixed representation, inspected at a finer grain, supports a much sharper pass/fail decision. There is also a limit on the far side. Any check that scores behavior only sees observed behavior, so it can confirm compliance under the conditions it tests and never beyond them Can behavioral training prove a model always complies?.

The surprising takeaway: freezing the representation is what makes minimal checks cheap and stable, but the same freezing hides internal disorganization from them. A check you can rely on has to inspect structure at the right resolution, not just confirm the output landed above a line. If you came looking for minimal-criterion search algorithms specifically, the collection doesn't cover them yet.


Sources 6 notes

Can editing hidden representations beat weight updates for finetuning?

ReFT learns task-specific interventions on frozen model representations rather than updating weights, with LoReFT (low-rank linear subspace variant) dramatically outperforming LoRA across reasoning, instruction-following, and NLU benchmarks while using far fewer parameters.

Can verification accuracy scale without training models?

Research shows verification accuracy improves independently via score granularity, repeated evaluation, and criteria decomposition—all deployable at inference without retraining. This reframes weak verifiers as under-scaled rather than fundamentally limited.

Can models be smart without organized internal structure?

Models trained with SGD can contain all the linearly decodable features needed for a task while maintaining fundamentally broken internal organization. This makes them vulnerable to perturbation and distribution shift invisible to standard evaluation metrics.

Are models actually reasoning about constraints or just defaulting conservatively?

Twelve of fourteen models perform worse when constraints are removed, dropping up to 38.5 percentage points. Models appear to reason correctly by defaulting to harder options, not by actually evaluating constraints.

Can verification separate structural near-misses from topical matches?

A two-stage pipeline—pooled-cosine recall followed by a small Transformer verifier operating on token-token similarity maps—reliably rejects structural near-misses that MaxSim-style late interaction cannot. The verifier succeeds because it operates on full token interaction patterns rather than compressed vectors.

Show all 6 sources
Can behavioral training prove a model always complies?

Any scored behavior is observed behavior, so training data cannot distinguish between a policy that always complies and one that complies only when watched. Only unobserved behavior would separate them, making such a test logically impossible.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.