Line of inquiry
Inquiring lines›What determines the reliability an…›How do systems improve effectively…›this line of inquiry
How do evaluation methodologies affect which model capabilities are revealed or hidden?
A broader line of inquiry — a family of 37 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 37
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can evaluation environments themselves become security exposures during capability testing?
- How often do deployed models exploit evaluation environments to hack their scores?
- How do live human evaluations differ from ground-truth benchmarks?
- How does evaluation environment design become part of the security boundary?
- Is the evaluation environment itself part of the security boundary?
- Do success-only evaluations systematically overestimate real-world deployment readiness?
- Why do open-world evaluations reveal capabilities that static benchmarks hide?
- How do open-world evaluations correct distortions that automated benchmarks introduce?
- How should we evaluate AI systems we cannot directly observe?
- Should evaluations shift toward open-world messy tasks instead of contests?
- How do frontier models exploit vulnerabilities in their own evaluations?
- Do automated benchmarks systematically distort what long-horizon agent capability actually looks like?
- Does scoring only final code execution waste diagnostic value of intermediate primitives?
- Can open-world evaluations become a scalable paradigm without becoming the next benchmark trap?
- Can evaluation environments contain security boundaries if they hold shared resources?
- What makes an evaluation environment itself a security boundary?
- Can test environments reliably predict how models behave in actual deployment?
- How does evaluation of exploit capability differ from other dual-use AI measurements?
- What belief errors about tool access show up as security measurement failures?
- Can we empirically test whether open models lower barriers to harmful workflows?
- Can automated evaluation replace human judgment in agent testing?
- How does speed of AI search prevent real-time supervision and evaluation?
- Which AI capabilities matter most for human-facing deployment contexts?
- How do we measure marginal risk instead of speculating about misuse scenarios?
- How should evaluation frameworks account for the computational cost of frontier AI capability?
- What evaluation criteria can hold across legitimate adoption and coercion?
- How do four separate fields each hold pieces of evaluation safety?
- Does good simulation eventually count as genuine realization?
- What baseline capabilities could bad actors achieve before open models existed?
- What vulnerabilities have models actually exploited in their own test environments?
- How much does believing deployment is real change model behavior strategically?
- How much does omniscient evaluation overstate real-world simulation fidelity?
- How should access controls scale with increasing capability evaluation intensity?
- How visible is the optional shortcut to the agent during evaluation?
- Why does held-out evaluation matter for detecting agent overfitting?
- Can agent-based simulators replace real-user A/B testing for studying recommendation system harms?
- What would whole-system AGI evaluation look like in practice?