Line of inquiry
Inquiring lines›How do we evaluate and improve AI…›What factors determine agentic sys…›this line of inquiry
What should agent evaluation prioritize to reveal reliable behavior?
A broader line of inquiry — a family of 47 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 47
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Which interaction artifacts matter most for reliable agent evaluation?
- Should feedback channels be excluded from the reward path in agent evaluations?
- Why do outcome-level metrics fail to reveal contained attacks in multi-agent pipelines?
- What makes some agent benchmarks measure interaction quality better than others?
- How do agent privacy compliance and task success differ in evaluation?
- Can a reward-seeking agent be distinguished from one pursuing intended behavior?
- Should artifact-level benchmarks replace token counts for agent evaluation?
- Does shared experimental state alone explain progress or is an analyzer needed?
- Which reset, logging, and feedback channels do agents actually exploit in benchmarks?
- Can measures of application actions reveal changes in coordination that output metrics miss?
- Can an agent change reward-path state through actions during evaluation?
- Can automated evaluation replace human judgment in agent testing?
- Does scoring only final code execution waste diagnostic value of intermediate primitives?
- What infrastructure and reporting standards would make interactive evaluation reproducible?
- How do agent actions change state that reward procedures later read?
- Can per-stage results reveal which demand causes agent failure in exploitation?
- Why do scalar evaluation scores collapse distinguishable agent behaviors?
- What makes a win untrustworthy in hidden evaluation environments?
- Why do current evaluation metrics fail to catch reasoning failures in persona agents?
- What makes a correct scoring function report misleading results in agent evaluations?
- How much of an agent's behavior actually escapes human review in practice?
- How do evaluation methods differ for single versus multi-agent systems?
- What other agent behaviors besides citations reveal reasoning quality?
- What does agent security look like when measured across interaction trajectories?
- Can task success alone reveal whether memory routing is working?
- What makes exploration and reflection rewards verifiable in agentic environments?
- Why does moving the reward target prevent saturation better than finding a better static proxy?
- Why does decoupling evaluation into components make hacking more diagnosable?
- How do evaluation hacks differ from genuine sandbox escapes?
- What governance and safety measurements matter for deployed agent environments?
- What within-run behavioral dimensions reveal where long-horizon agents succeed or fail?
- What counts as research completeness versus correctness in agent evaluation?
- Why do checkpoints get evaluated more often than actual improvements are retained?
- How do minimal-disclosure privacy contracts enable multi-dimensional agent evaluation?
- Why does held-out evaluation matter for detecting agent overfitting?
- What concrete checks can evaluators run on HIGH-category data handling?
- What validates whether a rewritten agent is actually better?
- How do you verify agent code under incomplete feedback signals?
- What role does runtime feedback play in agent verification and progress confirmation?
- Can gradients extracted with agent-level supervision transfer across different benchmarks?
- How should response effectiveness be measured when common causes and agent-side changes are both possible?
- How do you find which actions belong together before evaluation?
- What cognitive capabilities do agents need to internalize social feedback?
- Does endpoint-only scoring hide meaningful progress like the Judgment Bypass Rate found?
- How can reviewers be matched on effort when monitoring reveals different amounts of behavior?
- How does execution-guided critique differ from abstract action evaluation?
- What makes idle window detection valuable for continuous agent improvement?