Line of inquiry
Inquiring lines›What drives capability improvement…›How should computational architect…›this line of inquiry
What explains the gap between benchmark scores and true reasoning capability?
A broader line of inquiry — a family of 90 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 90
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do benchmark scores rise while reasoning quality declines?
- Do reasoning benchmarks predict real performance in long delegated workflows?
- Can benchmark scores on verifiable tasks transfer to unseen problems outside the training domain?
- Why do estimates for task-level performance differ so much from full job automation timelines?
- How does measurement error in capability benchmarks systematically underestimate or overestimate true ability?
- Can identical model performance mask fundamentally broken internal representations?
- What is the gap between benchmark performance and real workplace task completion?
- How can benchmark accuracy scores mask the absence of interpretable reasoning structure?
- How often do metric improvements fail to reflect real capability gains?
- Can memorization inflate apparent capability on benchmarks with available solutions?
- Why do internal representations differ when external performance matches?
- Should benchmark evaluations use multiple prompt formulations for difficult tasks?
- Can a model be strong at MMLU but weak at long-horizon tasks?
- How does optimizing model performance decouple from optimizing user interpretability?
- Why do text-only benchmarks underestimate deployed model capability?
- What real-world tasks most clearly expose gaps between benchmark performance and actual capability?
- Do perfect accuracy scores hide broken internal representations?
- Why do benchmark tests fail to detect LLM comprehension gaps?
- Why does a rising score not always mean improving capability?
- Why do task-completion benchmarks miss the competence of knowing when to abstain?
- How does a single score mix exploitation ability with task capability?
- How do frontier models maintain agreement scores above 90 percent across reasoning tasks?
- What evaluation methods actually measure reasoning versus execution capability?
- What drives the mismatch between general benchmark leadership and task-specific performance?
- Do current math benchmarks measure outcomes or rhetorical plausibility?
- How does tool access change what we measure in reasoning tests?
- How can interactive evaluation avoid replicating fragmentation problems from response-centered benchmark culture?
- What distortions do automated benchmarks introduce compared to real tasks?
- How do automated evaluation metrics differ from human expert judgment?
- Why do short interaction benchmarks fail to predict long horizon performance?
- Can high test performance mask a complete absence of understanding?
- How do static benchmarks fail to capture human preference alignment?
- What biases affect how we measure progress on research leaderboards?
- Why do benchmarks become saturated so quickly after initial launch?
- How do weight perturbations reveal what performance benchmarks cannot measure?
- Why do majority-label benchmarks hide models' failure on subjective tasks?
- What other hidden biases might aggregate metrics fail to distinguish from reasoning?
- Can clean benchmarks reveal true RLVR reasoning gains?
- Do alignment benchmarks measure actual bias removal or only verbal compliance?
- Can the benchmark-performance mismatch be estimated reliably without an oracle?
- Do short interaction benchmarks predict how LLMs perform in long workflows?
- Can automated benchmarks fairly evaluate messy real-world research tasks?
- How much of weak-to-strong performance gaps reflect presentation rather than capability?
- How much RLVR improvement comes from benchmark data memorization?
- How should benchmarks test whether models fit algorithms or patterns?
- What demographic and source limitations affect representativeness of expert-written benchmark datasets?
- How does benchmark performance measure translate to general self-modification ability?
- Does monitor position in the optimization loop matter more than capability gaps?
- How much do metric choices inflate claims about model capabilities?
- What distinguishes genuine task improvement from evaluator exploitation?
- Can an average-case validator score hide poor performance on critical tasks?
- Does measured performance gain reflect true task improvement or evaluator exploitation?
- How should benchmarks evaluate workflow architecture versus raw model performance?
- Which benchmarks benefit most from adding a separate memory module?
- Does the Heuristic Override Benchmark measure enumeration or world knowledge?
- How much improvement comes from caching versus actual capability gain?
- Does semantic auditing of instruction data improve performance uniformly across different model sizes?
- Can standard accuracy metrics miss the real constraints on user consumption?
- What language capabilities does fluency on standard benchmarks actually measure?
- Can empirical validation sustain long-term optimization without becoming gamed?
- Why do hidden test partitions matter more than open evaluation sets?
- Why might larger models become less honest despite better truthfulness scores?
- What deployment context determines which benchmark mode actually matters?
- Are cheap testbeds and skewed task distributions linked by design necessity?
- Why do leaderboard metrics fail to capture human flourishing in LLM evaluation?
- How should we redesign benchmarks to catch conservative bias in reasoning tasks?
- Can benchmarks designed for shortcut learning detect heuristic override failures?
- Why do current benchmarks fail to match user satisfaction with search results?
- Why do speech benchmarks still measure transcription instead of comprehension?
- Why do single function-calling benchmarks mask model weakness in specific areas?
- Can static analysis derive task bindings without manual effort?
- How should single-axis benchmarks account for separable capability dimensions?
- Should long-context evaluation measure the coupled system?
- Can a single Elo ranking represent multidimensional model capability?
- What role does vague intent play in realistic search evaluation?
- Are larger models and search access substitutes for factual accuracy?
- Should long horizon performance be measured as a separate evaluation axis?
- Which specific AI R&D tasks does AIDE2 benchmark itself against during selection?
- Can text-space optimization and audit governance coexist in a single skill lifecycle?
- How do AIDE2's held-out gains compare to matched-budget test-time search baselines?
- Can a single competence score capture multiple separable dimensions of capability?
- How much of MATH-500 improvement comes from data contamination versus real reasoning gains?
- Can contextual design decisions resist formalization into evaluation rubrics?
- How much do prompts and data splits shift a single benchmark score?
- How should query augmentation strategies be properly evaluated against baselines?
- How does task contamination differ from test set data leakage?
- How widespread is task contamination in LLM evaluation benchmarks today?
- How does score granularity connect to verification as a scaling axis?
- Who validates task bindings and how is validation checked?
- What fraction of the paper's tasks were actually misspecified or easy to hack?