Line of inquiry
Inquiring lines›Why are language models fragile de…›How do learned model representatio…›this line of inquiry
How do benchmark design choices systematically hide LLM limitations?
A broader line of inquiry — a family of 34 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 34
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do standard language benchmarks underestimate what LLMs can actually do?
- Why do NLP benchmarks exclude ambiguous instances from evaluation?
- Why do NLP benchmarks systematically exclude ambiguous test cases from evaluation?
- Why do standard NLP benchmarks hide the most critical language limitations?
- Why do NLP benchmarks hide LLM failures in ambiguity handling?
- Why do benchmark scores rise while reasoning quality declines?
- Why do benchmark tests fail to detect LLM comprehension gaps?
- Why do NLP benchmarks treat annotation disagreement as noise rather than signal?
- How do static benchmarks fail to capture human preference alignment?
- How do human annotators disagree systematically on ambiguous examples?
- How do general language model benchmarks predict specialized domain performance?
- Why do backward-looking benchmarks underestimate LLM scientific value?
- Can a model be strong at MMLU but weak at long-horizon tasks?
- Why do current language model judges collapse into coarse discrete scores?
- How do frontier models maintain agreement scores above 90 percent across reasoning tasks?
- Do current math benchmarks measure outcomes or rhetorical plausibility?
- Can high test performance mask a complete absence of understanding?
- Do alignment benchmarks measure actual bias removal or only verbal compliance?
- Why are ground truth labels missing from unlabeled domain evaluations?
- Why do leaderboard metrics fail to capture human flourishing in LLM evaluation?
- Can an LLM be well calibrated but still unreliable on single evaluations?
- Can graded relevance assumptions hold when user ratings are temporally inconsistent?
- What language capabilities does fluency on standard benchmarks actually measure?
- Why do high-disagreement tasks benefit from broad rater pools over deep annotation?
- Where should measurement systems sit to avoid recording bias?
- How widespread is task contamination in LLM evaluation benchmarks today?
- What measurement artifacts emerge when annotators interpret the same question differently?
- Why do older datasets show higher LLM performance than newer ones?
- Why does a domain-conditional bound fail outside its calibrated workload?
- Do high-disagreement items signal contested values or measurement noise?
- Can similar outputs from different systems prove they work the same way?
- How does saturation-aware aggregation encourage balanced improvements across multiple rubric dimensions?
- How do hidden partitions in evaluators compare across training and selection substrates?
- Why do Llama-based models outperform GPT-4 in objective clinical guidance?