Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How do architectural choices affec…›this line of inquiry
What external process records should verify agent behavior and benchmark claims?
A broader line of inquiry — a family of 58 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 58
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why does infrastructure-side evidence matter more than agent-reported traces?
- Can infrastructure evidence ground benchmark claims better than terminal scores alone?
- What makes recorded transitions more trustworthy than agent reasoning trajectories?
- How should verifiable process memory anchor safety-critical action logs?
- What process records would independently verify that agents performed required steps?
- Can agents themselves read and rely on tamper-evident process records?
- What evidence should benchmark operators attach to completion claims?
- Can pinned artifacts prevent audit agents from making inconsistent judgments?
- Who holds authority to anchor evidence in this system?
- How does recording state provenance help detect unauthorized tampering between agent actions?
- Where should authenticated provenance records sit to remain outside agent reach?
- Can an auditor verify environment state without trusting the executor's self-report?
- How can operators ground benchmark completion claims in infrastructure data?
- What architectural controls secure capture authenticity beyond signing?
- How can anchored records fail authenticity while passing integrity checks?
- How do signed logs compare to externally anchored records for audit?
- When does an agent's action earlier in the loop change what a scorer reads later?
- Does the recorder producing evaluation evidence sit inside the security boundary?
- What makes a detector's output count as integrity evidence?
- What fraction of conditional-compliance reports come from agentic versus non-agentic settings?
- Who bears responsibility for reconstructing evidence when parties may be adversarial?
- Does inspectable skill artifacts guarantee the behavior matches the person it claims to ground?
- Do infrastructure event records alone suffice to distinguish different failure mechanisms?
- How should journals reward verification work if it becomes separated from discovery?
- How do execution traces and tests represent agent environment state?
- Who decides whether an entity has authority to anchor a record?
- How should researchers separate factual claims from systems lessons in preliminary incident reviews?
- What makes a claim robust enough to lift out of preliminary incident evidence?
- Can commitments prove the right content was captured, not just that it matches later?
- What does a verification verdict miss when required steps never run?
- Can protocol compliance certify that a validator's objectives remain aligned?
- How do forged approvals fail differently at token versus policy checks?
- Does replay fidelity hold when policies explore branches history never visited?
- Can missing recorded stops tell us whether mechanisms actually exist?
- What tests would reveal whether recorded human approvals represent real oversight?
- How do interpretive and evaluative disagreement show up differently in agent traces?
- Does the paper treat storage traces as addressed messages or unmarked traces?
- How should memory poisoning success be scored at the validator stage?
- What makes provenance infrastructure more critical than artifact quality?
- What does trajectory audit reveal about evolution cycle contributions and costs?
- What cost metrics does the paper report for each authorization component?
- Which specific EU AI Act provisions does anchored evidence satisfy or address?
- Does a blockchain anchor prevent tampering or only reveal it?
- Can a blockchain anchor distinguish when an event happened from when it was recorded?
- Does endpoint-only scoring hide meaningful progress like the Judgment Bypass Rate found?
- How do append-only Git records compare to continuously updated world models for research continuity?
- What access requirements limit interventional audits to white-box settings?
- What makes a process for choosing between values legitimate and fair?
- How does execution-guided critique differ from abstract action evaluation?
- Why do credentials need evidence standards beyond permission categories?
- Can merged pull requests serve as a meaningful quality measure?
- How do you handle disagreement between two accounts of the same incident?
- What triggered OpenAI to adopt transparency over delayed investigation in misalignment reporting?
- What financial incentives shape the Foundation's interpretation of this data?
- Where else in the vault are recovery and rollback mechanisms already specified?
- What training process caused Gemini 3.1 Pro's record-tampering behavior?
- How does rubber-stamping differ from loss of scrutiny capacity in review processes?
- Does adjudicated mean labels agree with each other or with some standard?