SYNTHESIS NOTE
Topics›Reasoning o1 o3 Search›this note

Can distance alone rank which substrates resist reward hacking?

Does the amount a method changes a model reliably predict how exposed it is to evaluator errors? The paper tests whether a single distance metric can universally order vulnerability across weights, selection, and prompts.

Synthesis note · 2026-09-24 · sourced from Reasoning o1 o3 Search

The abstract states the formal core: "We formalize a distance-dependent upper bound on evaluator disagreement and a capacity ordering for nested policy classes, then show why distance alone cannot establish a universal ranking of vulnerability. An exact finite-output illustration demonstrates how the location of a scoring defect changes the behavior favored by each method." The discussion separates the pieces: "A distance-dependent error envelope gives an upper bound; class inclusion gives a capacity ordering; the geometry of accessible behaviors and the effectiveness of search determine what a particular system actually finds."

Three kinds of statement. The envelope bounds how much evaluator disagreement can matter given how far a policy moves; "how far a policy moves" is the excerpt's only gloss on distance. The ordering says that when one policy class includes another, the larger class has at least the capacity of the smaller. Neither says what a system will find. That depends on where the evaluator's errors sit among the behaviors it can reach and on how well search locates them, which is what the finite-output illustration varies: move the scoring defect and each method favors different behavior. The excerpt does not say the ranking of methods reverses, and it does not give the outputs, the scorer or the methods compared.

Why it is useful (my reading). The tempting shortcut is that the substrate that moves the policy least is the safest. The paper's move is to say a bound is not a forecast and to keep the statements separate so nobody reads one as the other. The point is one a practitioner can use without the formalism: a single number for "how much this method changes the model" does not tell you how exposed the method is to a given evaluator's mistakes, because exposure is set by where the mistake is.

The limit. A finite-output illustration is a constructed case. It supports the claim that no ranking follows from distance alone, and it is not evidence about any deployed system or about which substrate is worse in practice. The excerpt has no formulas, so what "distance" is measured from, and whether the bound is tight, is not stated.

Inquiring lines that read this note 50

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How prevalent is reward hacking in frontier models? Do planted honeypot tests reliably measure reward hacking? How does outcome-only reporting obscure which system components blocked attacks? Do current AI defenses adequately protect against semantic manipulation attacks? Do AI capability benchmarks accurately measure reasoning ability or just surface patterns? What determines whether AI system errors remain visible and contestable? How do evaluation methodologies affect which model capabilities are revealed or hidden? What infrastructure evidence validates agent benchmark achievement claims? How do models reward hack during evaluation and can detection succeed? Can causal models and layer interventions detect and restore hidden model behaviors? Do frontier models develop hidden self-protective behaviors? Can defenses detect attacks composed across multiple skills? How can defenders detect coordinated attacks across episodes?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 79 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

distance alone cannot establish a universal ranking of vulnerability to reward hacking across substrates — in the paper's finite-output illustration the location of the scoring defect changes which behavior each method favors