Line of inquiry
Inquiring lines›What explains language model reaso…›How do contextual factors, biases,…›this line of inquiry
How do LLM judges' systematic biases affect alignment and evaluation outcomes?
A broader line of inquiry — a family of 44 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 44
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- What other evaluation biases exist in LLM judge systems?
- Can an LLM judge's bias be reduced through prompting or other interventions?
- How do LLM judges' built-in biases influence the policies they help align?
- Does LLM judge preference for LLM arguments amplify errors in contested factual domains?
- Do LLM judges with diverse personas resist individual biases better than single evaluators?
- Do LLM judges systematically favor arguments from other LLMs?
- Can parallel evaluation reduce position and length bias in LLM judging?
- What biases do single large LLM judges introduce into comparisons?
- Why do LLM judges systematically favor outputs from their own model family?
- Does ensembling smaller judges reduce bias more effectively than single large judges?
- Why do LLM judges show more extreme sycophancy bias than humans?
- How does LLM judge bias amplify errors in multi-agent debate on contested factual questions?
- What calibration corrections can reduce LLM judge bias in automated evaluation pipelines?
- What are the four catalogued biases that make LLM judges vulnerable to prompt attacks?
- What systematic biases do LLM judges introduce into AI-evaluated debates?
- Do smaller LLM judge panels outperform single large judges in practice?
- Why do current language model judges collapse into coarse discrete scores?
- Why do LLMs show gender bias but humans evaluators do not?
- Do language models inherit gender bias from training data in grading tasks?
- Does longer context and multi-step reasoning compound small preferences in graders?
- Which biases in LLM judges are exploitable through presentation alone?
- Why do LLMs systematically prefer text from their own family?
- What four exploitable biases make current LLM judges vulnerable to zero-shot attacks?
- Can LLMs evaluate logical argument quality in debates they themselves can win?
- Can crowdsourced voting and automated panels both credibly evaluate LLM outputs?
- What surface features do LLMs rely on when judging response quality?
- Why is consistency between argument scoring and winner selection lower for LLM judges?
- How robust is misalignment classification across different judge models?
- How sensitive are LLM bias measurements to analysis choices?
- Can instruction prompts reliably steer an LLM judge toward specific alignment targets?
- Can an LLM judge reliably report its own biases rather than remove them?
- Why do LLM judges assign high argument strength scores yet pick LLM winners anyway?
- Can LLM judges be trained to think more rigorously during evaluation?
- What shared epistemic faults persist even when judges come from different families?
- What biases might an LLM judge introduce into an on-policy alignment process?
- Why do LLMs swing on minor rewording yet ignore explicit bias correction instructions?
- How do calibration and reliability differ in LLM judge evaluations?
- Why do review corpora contain biases that affect generated comparisons?
- Can masking company identity in grading materials eliminate the bias?
- Can LLM judges reliably estimate when they lack sufficient persona information?
- Where should measurement systems sit to avoid recording bias?
- What role should stakeholders play in evaluating LLM fairness?
- How much better is a panel of smaller judges than one large judge?
- What does McDonald's omega reveal about LLM judgment consistency?