Line of inquiry
Inquiring lines›What makes reasoning better — more…›How do prompts and framing affect…›this line of inquiry
How do evaluation biases undermine LLM quality assessment systems?
A broader line of inquiry — a family of 32 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 32
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can LLMs reliably assess the quality of ideas they generate?
- Can language models accurately evaluate the quality of their own ideas?
- Does LLM judge preference for LLM arguments amplify errors in contested factual domains?
- Why do LLM outputs match researcher priors without solving tasks correctly?
- Can parallel evaluation reduce position and length bias in LLM judging?
- What other evaluation biases exist in LLM judge systems?
- Can researchers prevent their expectations from shaping LLM outputs?
- Why do backward-looking benchmarks underestimate LLM scientific value?
- Can human researchers improve LLM ideas through iterative feedback?
- Why does automated evaluation consistently overestimate research quality?
- Can crowdsourced voting and automated panels both credibly evaluate LLM outputs?
- Can structured decomposition fix evaluation gaps in other research tasks?
- Do LLM judges with diverse personas resist individual biases better than single evaluators?
- What capability boundary exists in LLM prediction of effect sizes?
- Why do leaderboard metrics fail to capture human flourishing in LLM evaluation?
- Why does probability of text completion not equal knowledge value?
- Does exposure to more domain-specific examples reduce LLM overconfidence?
- Can LLMs evaluate their own observations without external feedback?
- How do years of A/B testing compare to one-shot LLM content generation?
- Can LLMs learn to signal evaluative commitment through metadiscursive language?
- How do LLMs generate false citations that sound like real scholarship?
- Can LLM persuasion be fairly evaluated without stratifying by reader background?
- Why do some LLM clusters cite broader psychology than others?
- How does the absence of evaluative stance appear in LLM academic writing?
- Can proxy evaluation of ideas accurately predict their quality without implementation?
- How should moderator LLMs decide which speakers to query per topic?
- Why do LLM judges assign high argument strength scores yet pick LLM winners anyway?
- How does automated transcript analysis compare to patient self-report on engagement?
- Why does LLM fluency create false perceptions of professional standing and expertise?
- How does semantic entropy compare to confidence scores from internal model probabilities?
- Can knowledge density explain why LLM writing feels coherent but fatiguing?
- What does McDonald's omega reveal about LLM judgment consistency?