Line of inquiry
Inquiring lines›What explains language model reaso…›Why do models produce unreliable r…›this line of inquiry
How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking?
A broader line of inquiry — a family of 40 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 40
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can judge bias be contained by system design rather than prompted away?
- Can critic model trios evaluate reasoning quality more reliably than outcome rewards alone?
- Does debate training prevent reward hacking when judges show preference bias?
- Can judges trained on both verifiable and non-verifiable tasks transfer across domains?
- Can counterfactual invariance techniques address exploitable biases in LLM judges?
- Does meta-judging improve evaluator quality better than temporal decoupling alone?
- Do situationally aware models deliberately exploit their graders' judgment gaps?
- Can synthetic disagreement tests reliably measure hidden reward-seeking?
- Can a static evaluator become the performance ceiling for an improving actor?
- How do ensemble methods reduce bias in automated evaluation?
- Why does a relativistic critic outperform absolute scoring in adversarial reasoning training?
- Can evaluators detect value-driven output biases without comparing paired questions?
- Can verifier output replace ground-truth answers as the asymmetric information source?
- Can systems recognize and abstain on judgments rather than hallucinating preferences?
- Can proper scoring rules fix RLVR's degradation on disagreement prediction?
- Can adversarial critics force genuine reasoning the same way critique fine-tuning does?
- Can environment feedback alone provide dense credit without a teacher?
- Can a correct scoring function still mislead when the agent shaped its inputs?
- How might automated evals eventually capture the human judgment designers exercise now?
- What happens when LLMs grade other LLMs in closed evaluation loops?
- How does same-author bias interact with the four adversarial judge biases already documented?
- How can judges evaluate thinking without seeing the actual thoughts?
- How much of observed stance reversal actually harms user decision-making in practice?
- How does comparing answers differ from answering when activating company preference?
- Do own-company biases differ across model families in grading tasks?
- Why do human raters miss factual errors that domain experts catch?
- Does disjoint family diversity actually cancel model-specific bias in evaluation?
- Can fact-checking labels replace the cultural work of developing a discount?
- Can a diverse panel approach work for validators beyond text evaluation?
- Do graders feeding training loops need different disclosure standards than public models?
- Can proxy evaluation of ideas accurately predict their quality without implementation?
- Why are expensive rankers more resilient to adversarial content than cheap ones?
- How does information asymmetry between teacher and student create the learning signal?
- Why does strengthening the judge improve the actor's generation performance?
- When do aggregated imperfect demonstrations fail to outperform the best expert?
- Why do static evaluators become a constraint on model improvement over time?
- Why does masking future experts guarantee causal validity without external verification?
- Why does information asymmetry between teacher and student enable effective feedback learning?
- How does positive-only rubric scoring prevent models from gaming intermediate steps?
- Can unified policies handle negative feedback and critique transformation simultaneously?