Line of inquiry
Inquiring lines›How do agents behave and coordinat…›How can we measure and understand…›this line of inquiry
Why do agents report success when they have actually failed?
A broader line of inquiry — a family of 37 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 37
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do agents report success when actions actually fail?
- Why do agents report success when their actions actually fail?
- How often do agents report success when their actions actually failed?
- Why do autonomous agents report success on failed actions?
- Why do agents report success when they have actually failed at tasks?
- How do agents learn to report success on actions that actually failed?
- Can confident agent failures appear as successes in outcome reporting systems?
- Which failure modes dominate in autonomous research agents?
- Why do confident failures on failed actions become a signature problem?
- How does completion bias in agents differ from other epistemic failure modes?
- What tasks do AI agents still fail at most often?
- Why do AI agents fail at verification but succeed at generation?
- How do agents decide when to stop and reflect on failure?
- Can agent success reports serve as reliable oversight signals in real deployment?
- Why does human interaction remain the hardest failure mode for agents?
- What happens when an agent judges its task impossible?
- How should safety systems catch confident failures from agents that report success on unsafe actions?
- Why do agents make premature commitments when user goals are still forming?
- How much autonomy can agents safely exercise before failing?
- How do mode-specific failures differ between completion and agent benchmarks?
- What distinguishes mechanical generation failures from deliberate behavioral withholding?
- Why do completion-mode strengths not transfer to agentic settings?
- What causes the gap between agent reasoning and agent action?
- Can semantic audit layers attribute failure mechanisms to infrastructure-level state changes?
- What structural features enable agents to detect when understanding has broken down?
- What makes action-producing models fail in ways text models typically do not?
- What are the differences between chat model and agent authorization failures?
- How does poor belief tracking cause agents to keep acting past the point of usefulness?
- What distinguishes confident failure from deliberate alignment faking in agent behavior?
- What training objectives could reduce completion bias in autonomous agents?
- Can the same test failure come from incentive problems versus information failures?
- How do delayed effects complicate causal attribution in agent systems?
- Does an agent stop work or escalate when it cannot complete an assigned task?
- What are the fourteen failure modes in deep research agents?
- What distinguishes honest Byzantine faults from epistemic faults?
- Why do LLM agents make promises without executing them?
- What distinguishes first-order from second-order agency in language models?