Line of inquiry
Inquiring lines›How do agents behave and coordinat…›How do agents in production pipeli…›this line of inquiry
How can we verify agent claims against their actual capabilities and actions?
A broader line of inquiry — a family of 46 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 46
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do agents systematically misreport their own capabilities and tool access?
- Which interaction artifacts matter most for reliable agent evaluation?
- What makes recorded transitions more trustworthy than agent reasoning trajectories?
- How should harness infrastructure validate code that agents generate themselves?
- Do agents interpret peer edits as legitimate prior changes versus tampering?
- How should we label ground truth when a protected state change alone is ambiguous?
- Why does forcing agents to trace function paths prevent unsupported claims?
- Can agents themselves read and rely on tamper-evident process records?
- How can verifiers check policy compliance in agentic reasoning tasks?
- What process records would independently verify that agents performed required steps?
- Does held-out validation prevent skill document edits from drifting or accumulating harm?
- What role does peer activity play in triggering protected test modifications?
- How should human oversight apply to persistent agent-authored code?
- How do agent actions change state that reward procedures later read?
- How reliable is agent self-description compared to infrastructure monitoring for detecting intent?
- What other agent behaviors besides citations reveal reasoning quality?
- How should verifiable process memory anchor safety-critical action logs?
- Can agents rationalize rule violations by reframing them as repairs?
- How much of an agent's behavior actually escapes human review in practice?
- What makes exploration and reflection rewards verifiable in agentic environments?
- Who holds authority to anchor evidence in this system?
- What architectural controls secure capture authenticity beyond signing?
- Why can agent-restored files pass correct checks but violate task intent?
- When does an agent's action earlier in the loop change what a scorer reads later?
- How does recording state provenance help detect unauthorized tampering between agent actions?
- What must auditors reconstruct when reviewing an agentic workflow decision?
- Where should authenticated provenance records sit to remain outside agent reach?
- Can pinned artifacts prevent audit agents from making inconsistent judgments?
- Can an agent weaken a test or restore files to change what the grader checks?
- How can agents distinguish between optional and required form fields during execution?
- What concrete checks can evaluators run on HIGH-category data handling?
- Who bears responsibility for reconstructing evidence when parties may be adversarial?
- What validates whether a rewritten agent is actually better?
- Does accountability differ when one party in an exchange cannot hold commitments?
- Why does correcting an agent's objective leave its available actions unchanged?
- What role do false beliefs play in agents violating protected requirements?
- What specific training mechanism causes agents to over-claim actions and overwrite documents?
- Does AIDE2's guard against bad wins sit inside or outside the rewritable code?
- Can a policy distinguish genuine objects from traps without revealing that distinction?
- Does inspectable skill artifacts guarantee the behavior matches the person it claims to ground?
- How do capability tracks and behavior tracks stay separable during skill deployment?
- What signals reveal when agents first touch an artifact they did not create?
- Does an agent's own prior conduct shape the counterparty's response?
- Did AIDE2's rewrites solve problems on a human checklist or search artifacts?
- How do cognitive state traps compromise agent-writable monitoring history?
- Why do uncommitted changes create ambiguity about preserving versus restoring state?