Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›How can systems ensure safety and…›this line of inquiry
Do reasoning benchmarks predict model performance in long-horizon workflows?
A broader line of inquiry — a family of 51 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 51
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do reasoning benchmarks predict real performance in long delegated workflows?
- Should benchmarks measure trace length or whether constraints were actually satisfied?
- Why do estimates for task-level performance differ so much from full job automation timelines?
- Can a model be strong at MMLU but weak at long-horizon tasks?
- What is the gap between benchmark performance and real workplace task completion?
- Should benchmark evaluations use multiple prompt formulations for difficult tasks?
- How does optimizing model performance decouple from optimizing user interpretability?
- Can static analysis derive task bindings without manual effort?
- Why do sparse per-step errors accumulate undetected across delegated tasks?
- How does tool access change what we measure in reasoning tests?
- How does accumulated context history degrade iteration quality in long-horizon tasks?
- Does semantic auditing of instruction data improve performance uniformly across different model sizes?
- Do short interaction benchmarks predict how LLMs perform in long workflows?
- Why do NLP benchmarks systematically exclude ambiguous test cases from evaluation?
- Can automated benchmarks fairly evaluate messy real-world research tasks?
- Can verification cost be measured separately from task completion speed?
- Can clean benchmarks reveal true RLVR reasoning gains?
- What distinguishes genuine task improvement from evaluator exploitation?
- Can tool use or self-conditioning fix degradation in extended LLM workflows?
- Which model capabilities actually matter for sustained workflow delegation?
- Can structured output formats reduce instruction following degradation?
- Which benchmarks benefit most from adding a separate memory module?
- Does the Heuristic Override Benchmark measure enumeration or world knowledge?
- How should benchmark design account for task-dependent sparsity tolerance differences?
- Can text-space optimization and audit governance coexist in a single skill lifecycle?
- How should benchmarks evaluate workflow architecture versus raw model performance?
- Do distributed relational tasks consistently underperform local classification across NLP domains?
- Which specific AI R&D tasks does AIDE2 benchmark itself against during selection?
- How do memory relevance filters fail to prevent performance degradation?
- Should long-context evaluation measure the coupled system?
- Can benchmarks designed for shortcut learning detect heuristic override failures?
- Can dynamic evidence collection improve task verification accuracy?
- Do synthetic verification chains from long-CoT models match the quality of human-annotated process labels?
- Why does a domain-conditional bound fail outside its calibrated workload?
- Who validates task bindings and how is validation checked?
- How do AIDE2's held-out gains compare to matched-budget test-time search baselines?
- How are conflict tasks constructed to test alignment between model and user intent?
- How does task contamination differ from test set data leakage?
- Why do models that excel at task success often fail at privacy compliance?
- How are task bindings validated and what does validation cost per task?
- How widespread is task contamination in LLM evaluation benchmarks today?
- Can a complexity-predictor be meaningful if models are redundant?
- What fraction of the paper's tasks were actually misspecified or easy to hack?
- What makes some analysis tasks stable enough for rigid generated interfaces?
- Why does homework adherence remain low despite advances in language model capability?
- At what interaction length does MCP's application-layer state code become unwieldy?
- How many task-specific bindings does BenchShield require across benchmarks?
- What makes out-of-band monitoring better than in-band verification loops?
- How do execution traces represent state and dynamics in codebase modeling?
- What event types and phases structure the BenchShield lifecycle model?
- How were ten thousand scenarios validated across fifty domains?