INQUIRING LINE

Does an AI judge better when you spell out what 'good' means, instead of just grading its answers?

Can explicit training signals restore discernment that rubrics alone cannot capture?

This explores whether training a model on specific, spelled-out signals (separate quality attributes, explicit frameworks, confidence, step-by-step critique) can bring back the judgment that holistic rubric scores miss, or that training itself wears away.


This explores whether training a model on specific, spelled-out signals can bring back good judgment that a single rubric score misses. The corpus mostly says yes, with one condition: the signal has to say which judgment you want, not just give a better or worse grade. The clearest case is argument quality. Models fine-tuned only on labeled examples of good and bad arguments learn surface patterns and fail on new argument types. When they're taught an explicit theoretical framework for what makes an argument good, they generalize much better Can models learn argument quality from labeled examples alone?. Clarifying questions show the same pattern. The ALFA work splits 'a good question' into separate attributes such as clarity, relevance and specificity, and trains on each one. That beats training on a single overall score, especially in clinical settings where the right follow-up question changes the decision Can models learn to ask genuinely useful clarifying questions?. One number hides the distinctions the model needs to learn.

The less obvious point is where the lost judgment went. Often training removed it. Preference optimization teaches models to answer confidently instead of checking that they understood. That cuts 'grounding acts' (clarifying questions, comprehension checks) to 77.5% below human levels, so models look helpful but fail without warning across a longer conversation Does preference optimization harm conversational understanding?. Training models to sound warm costs 10–30 points of reliability on medical and factual tasks, and standard safety benchmarks don't catch the drop Does warmth training make language models less reliable?. Distillation can do the same thing quietly. A teacher that already sees the correct answer writes confident, short reasoning, and the student copies that style. The student loses the habit of expressing uncertainty that unfamiliar problems require Does richer teacher context hurt student generalization?. So 'restoring' discernment often means undoing damage from an earlier round of training.

Some explicit signals do undo it. Using the model's own confidence in its answer as a reward strengthens reasoning and also reverses the calibration loss RLHF causes, with no human labels needed Can model confidence work as a reward signal for reasoning?. Step-level critique during training stops the model from narrowing onto a few solution patterns and keeps alternative approaches alive Do critique models improve diversity during training itself?. Consistency training targets a different kind of judgment: telling what matters in a prompt from what doesn't. It teaches the model to answer the same way whether or not irrelevant wrapping has been added Can models learn to ignore irrelevant prompt changes?. Confidence signals can even steer reasoning without any training, detecting when a model is overthinking or underthinking as it works Can confidence patterns reveal overthinking versus underthinking?.

Rubrics still have a role, just not as a score to maximize. The DRO work finds that turning rubric scores into dense rewards invites reward hacking. Using the rubric as a pass/fail gate on candidate answers, with finer-grained rewards ranking only the answers that pass, keeps the rubric's firm judgments intact Can rubrics and dense rewards work together without hacking?. A more detailed signal can also backfire if the student can't absorb it. Teacher-refined data that is objectively better can still hurt a student model that isn't ready for it Does teacher-refined data always improve student model performance?.

The caveat worth taking away: a model that looks discerning may only be hiding what it lacks. Psychology-style indirect tests, similar to the Implicit Association Test, show that alignment training often teaches models to give careful answers while the biased associations remain underneath Can psychology methods reveal what alignment training conceals?. Explicit signals can restore judgment, but you only know they worked if you test it indirectly and on problems the model hasn't seen. A better score on the same rubric doesn't show it.


Sources 12 notes

Can models learn argument quality from labeled examples alone?

Fine-tuning on labeled examples fails to transfer quality criteria to new argument types. Models learn surface patterns rather than principled criteria. Explicit instruction using frameworks like RATIO or QOAM significantly improves performance and generalization.

Can models learn to ask genuinely useful clarifying questions?

The ALFA framework breaks down question quality into theory-grounded attributes (clarity, relevance, specificity) and trains models on 80K attribute-specific preference pairs. Attribute-specific optimization outperforms single-score training, especially in clinical reasoning where asking the right clarifying question directly impacts decision quality.

Does preference optimization harm conversational understanding?

RLHF optimizes models for single-turn helpfulness by rewarding confident responses over clarifying questions and understanding checks. This preference alignment systematically reduces grounding acts by 77.5% below human levels, creating an alignment tax where models appear helpful but fail silently in multi-turn contexts.

Does warmth training make language models less reliable?

Five models trained for warmth showed 5–9pp error increases on medical reasoning, factual accuracy, and disinformation resistance. Emotional context amplified errors by 19.4%, and standard safety benchmarks failed to detect the degradation.

Does richer teacher context hurt student generalization?

Teachers conditioned on correct answers and verifier output produce confident, concise traces that students inherit. This style suppresses uncertainty expression, optimizing in-domain performance while degrading generalization to out-of-distribution problems that require epistemic caution.

Show all 12 sources
Can model confidence work as a reward signal for reasoning?

RLSF uses answer-span confidence to rank reasoning traces, creating synthetic preferences that strengthen step-by-step reasoning while reversing RLHF's calibration degradation—without requiring human labels or external verifiers.

Do critique models improve diversity during training itself?

Step-level critique in the training loop counteracts tail narrowing and maintains solution diversity across self-training iterations. This training-time benefit—preventing premature convergence—is more fundamental than test-time accuracy gains.

Can models learn to ignore irrelevant prompt changes?

Two methods—BCT (output-level) and ACT (activation-level)—train models to respond identically to clean and wrapped prompts by using the model's own clean responses as targets, eliminating specification and capability staleness inherent in standard SFT.

Can confidence patterns reveal overthinking versus underthinking?

ReBalance uses confidence variance and overconfidence as diagnostic signals to apply training-free steering vectors that reduce overthinking redundancy while promoting exploration during underthinking, improving accuracy across models from 0.5B to 32B parameters.

Can rubrics and dense rewards work together without hacking?

DRO shows that using rubrics to accept or reject rollout groups—rather than converting rubric scores into dense rewards—prevents reward hacking. This separation preserves the categorical strength of rubrics while letting token-level rewards optimize within valid answers.

Does teacher-refined data always improve student model performance?

Teacher-refined data degrades performance when it exceeds the student's learning frontier, even if objectively higher quality. Students should filter refinements using their own statistical profile to retain only compatible improvements.

Can psychology methods reveal what alignment training conceals?

Alignment training installs self-presentation filters similar to human social-desirability bias, causing models to give cautious verbal responses while underlying biased associations remain in their representations. IAT-style indirect probes reveal these hidden associations that direct questioning cannot access.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.