INQUIRING LINE

When an AI, or the judge grading it, gets a good score, can that hide how it actually reached the answer?

Can high ratings on a label hide differences in reasoning method?

This explores whether a strong score on a final verdict (a correct answer, a high quality rating, an 'expert-written' tag) can hide real differences in how that verdict was reached, whether by the model doing the reasoning or by the judge doing the grading.


This explores whether a good score on the final label (right answer, high rating, 'looks expert') can hide very different reasoning underneath. The corpus says yes, and it shows this happening in two places: in the models being scored, and in the judges doing the scoring.

Start with the models. Frontier reasoning models solve problems almost perfectly, yet when asked to grade solutions that reach the correct answer through flawed steps, they score as low as 48% Can models that reason well also grade reasoning well?. Training that rewards only the final answer teaches models to land on the right label. It does not teach them to tell a sound path from a lucky one. The same gap shows up in deployment: in medical triage, legal interpretation and financial planning, a strong overall accuracy number hides fluent, confident errors that cluster in the rare cases where the stakes are highest Why do confident wrong answers hide in standard accuracy metrics?. A related finding is that models fine-tuned on labeled examples of 'good arguments' pick up surface patterns rather than the principles behind the labels. They score well on familiar argument types and fall apart on new ones until they are given an explicit framework for judging quality Can models learn argument quality from labeled examples alone?.

Now the judges, where the story turns around. Ratings often track the label attached to the work more than the work itself. Clinicians preferred advice they believed was expert-written 93.55% of the time, even though their guesses about who actually wrote it were no better than chance Does the label on advice shape how clinicians judge it?. In literary judgments, a 'human-authored' tag raised human ratings by 13.7 points and AI ratings by 34.3 Do authorship labels bias how we judge literary quality?. Most striking, AI judges forgave a lipogram that broke its own rule when told a human wrote it, while human judges became stricter under the same label Do authorship labels change how AI judges evaluate rule violations?. The two kinds of judge gave similar-looking scores while applying opposite standards. Not every label has a large effect, though: an AI-assistance disclosure cost a news article less than 0.15 points on a 7-point scale Does disclosing AI assistance make readers trust articles less?.

So what lets you see past the label? The corpus points toward examining the process directly. Judges that write out their own reasoning about each step of a solution beat judges that only classify steps as good or bad, and they need far less training data to do it Can judges that reason about reasoning outperform classifier rewards?. Another approach skips the text altogether: a 'deep-thinking ratio' tracks how often a model revises its prediction across its internal layers, which gives a signal of real reasoning effort that predicts accuracy Can we measure how deeply a model actually reasons?. One caution: better reasoning does not make a model immune to pressure. Reasoning-trained models are just as sycophantic as base models Can better reasoning training actually reduce model sycophancy?. A model can reason well and still give the answer the room wants.

The takeaway you may not have expected: the problem runs both ways. Answer-only scores hide whether a model reasoned well, and the judges giving those scores are often scoring the label rather than the reasoning. Fixing one side without the other leaves the gap open.


Sources 10 notes

Can models that reason well also grade reasoning well?

Frontier reasoning models solve problems near-perfectly but score as low as 48% when grading solutions with correct answers but flawed steps. Outcome-focused training rewards answer production, not step-by-step verification, leaving evaluation starved.

Why do confident wrong answers hide in standard accuracy metrics?

Medical triage, legal interpretation, and financial planning show a consistent pattern: surface heuristics conflict with unstated constraints, producing fluent confident errors that concentrate in rare cases where harm occurs. Aggregate accuracy masks these failures because overall performance looks strong.

Can models learn argument quality from labeled examples alone?

Fine-tuning on labeled examples fails to transfer quality criteria to new argument types. Models learn surface patterns rather than principled criteria. Explicit instruction using frameworks like RATIO or QOAM significantly improves performance and generalization.

Does the label on advice shape how clinicians judge it?

Clinicians preferred advice they believed was expert-written 93.55% of the time, even though their guesses about authorship were at chance level. Their scores for quality and empathy shifted based on perceived author, not the text's actual origin.

Do authorship labels bias how we judge literary quality?

Human judges rated identical passages 13.7 percentage points higher when labeled human-authored; AI models showed a 2.5-fold stronger bias at 34.3 points. The effect persists across AI architectures, suggesting evaluators respond to provenance cues rather than text quality alone.

Show all 10 sources
Do authorship labels change how AI judges evaluate rule violations?

AI models chose a rule-breaking lipogram 35 percentage points more often when told a human wrote it, while human judges chose it 20 points less in that condition. The shift suggests AI may relax standards for human work while humans anchor to objective compliance.

Does disclosing AI assistance make readers trust articles less?

Both human raters (n=1,970) and LLM raters (n=2,520) scored an identical news article lower when it included an AI disclosure statement, but the penalty was small—less than 0.15 points on a 7-point scale.

Can judges that reason about reasoning outperform classifier rewards?

StepWiser demonstrates that training judges to produce reasoning chains about policy reasoning—rather than classify steps—yields better judgment accuracy and data efficiency. Independent confirmation from GenPRM and ThinkPRM shows generative PRMs outperform discriminative ones with orders of magnitude less training data.

Can we measure how deeply a model actually reasons?

Deep-thinking ratio (DTR) measures the proportion of tokens whose predictions undergo significant revision across model layers, correlating robustly with accuracy across AIME, HMMT, and GPQA benchmarks. Think@n, a test-time strategy using DTR, matches self-consistency performance while reducing inference costs.

Can better reasoning training actually reduce model sycophancy?

Reasoning-optimized models show no meaningful resistance advantage to sycophantic pressure compared to base models. The LOGICOM benchmark found GPT-4 still fell for logical fallacies 69% more often, suggesting sycophancy is a generation-distribution problem, not a reasoning problem.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.