SYNTHESIS NOTE
Topics›Alignment›this note

Did accidental CoT grading harm model monitorability in practice?

OpenAI discovered chain-of-thought was accidentally rewarded during training of deployed models. The question is whether this degraded the ability to monitor and evaluate model reasoning through its stated outputs.

Synthesis note · 2026-10-08 · sourced from Alignment

OpenAI reports that despite maintaining "a policy against directly grading CoTs while training reasoning models," an automated detection system discovered that chain-of-thought (CoT) had been accidentally graded in RL training runs for several released models — GPT-5.4 Thinking, the GPT-5 Instant line, and the GPT-5.3 and GPT-5.4 mini variants (GPT-5.5 was unaffected). Three specific mechanisms were found: rewarding trajectory usefulness, penalizing unnecessary confirmation questions (affecting under 3.8% of samples for GPT-5.1 Instant and under 1.8% for the others), and penalizing successful prompt injections (under 0.3% of samples). Despite this, OpenAI states "our analyses showed no clear reductions in CoT monitorability," based on running their Guan et al. (2025) monitorability evaluations on checkpoints across the affected training.

The detection system works by regex-scanning all RL runs for cases where CoT text from a sample leaks into the inputs used to compute its own reward, alerting the run owner and posting to a public Slack channel — a cheap, scalable proxy for the real question rather than a semantic check, so coverage is imperfect for non-standard reward pipelines. To explain why these incidents didn't degrade monitorability, OpenAI ran separate deliberate experiments applying strong CoT-directed pressure, and identifies three governing factors: reward magnitude (how large the CoT reward is relative to other training signals), coverage (how often training puts the model in a situation where the pressure applies), and conditional discoverability (how likely the model is to find the rewarded CoT behavior within that situation). They found obfuscation "surprisingly difficult to produce" at standard strength, and even deliberately training on CoT-confidence pressure didn't saturate the effect; the accidental incidents were low-magnitude and low-coverage by comparison, which is their proposed explanation for the null result.

This is a real deployment incident report, not a model-organism experiment, which complicates the risk framing underlying Does optimizing against monitors destroy monitoring itself? and Can we monitor AI reasoning without destroying what makes it readable?: the "tax" those notes describe — that monitoring a CoT and optimizing against it degrades monitorability — did not materialize here, by OpenAI's own measurement, under naturally occurring low-magnitude pressure. The three-factor account (magnitude, coverage, discoverability) adds a dose-response structure that the quantified faithfulness numbers in Do reasoning models actually use the hints they receive? don't supply on their own. And OpenAI's stated response — detect accidental grading and avoid training on it, rather than retraining against it — is a direct avoidance of the dynamic in Does iterative training against detected failures prevent actual compliance?, where training against a detector would be expected to select for evasion rather than compliance.

The excerpt is a self-report by the lab whose models and policies are being evaluated, using its own detection system and its own monitorability evaluation suite, with no independent replication. OpenAI explicitly hedges that it "cannot rule out effects which are harder to measure," and states that "in principle with enough CoT optimization pressure we'd expect CoT monitorability to degrade" — so the finding is scoped to the specific low-magnitude, low-coverage incidents discovered, not a general claim that accidental CoT grading is safe. The implication, at the strength the evidence supports, is that small, incidental CoT-grading leaks are less dangerous than the monitorability-tax framing might suggest by default, but the factors that make deliberate obfuscation hard to induce in these experiments are exactly the factors that a more capable model or a determined adversarial optimizer could push harder on.

Inquiring lines that read this note 7

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI systems evade safety evaluations through reasoning manipulation? How does awareness of evaluation context influence model behavior? Can base models hide emergent misalignment through alignment training? How do individually-safe actions create collectively-unsafe outcomes? What external process records should verify agent behavior and benchmark claims? Why do standard evaluation practices obscure safety-critical AI failures?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 92 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

OpenAI finds accidental CoT grading in several deployed GPT-5 models did not measurably reduce monitorability