Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›How can systems ensure safety and…›this line of inquiry
How do evaluation practices shape which failures stay visible?
A broader line of inquiry — a family of 64 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 64
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How do response-centered evaluation assumptions hide safety-critical failure modes?
- Why do evaluation habits hide safety-critical challenges from view?
- What does it mean for errors to remain visible, contestable, and recoverable?
- Which evaluation habits keep safety-critical failures hidden in AI systems?
- Can AI outputs inspire new directions even when they seem like failures?
- How do inherited evaluation habits obscure failures that matter most?
- What conditions allow technical systems to escape critical evaluation?
- What makes a model's errors visible and contestable to users?
- What would it take to measure whether system errors stay visible and contestable?
- Does visibility and contestability of errors replace prevention as the safety goal?
- Why does greater automation actually obscure rather than eliminate research failure modes?
- Do frontier AI models fail in ways that preserve the appearance of competence?
- What makes the frame problem distinct from feature-level shortcuts?
- How does laboratory generalization evidence connect to deployment failure modes?
- How should learning environments balance error prevention with pedagogical value?
- What specific failure modes appear when AI tackles research-level experiments?
- What design principles prevent error cascades in multi-step evaluation systems?
- How do workflows normalize and hide errors before they become visible hazards?
- Does refining around bad results risk cascading errors in automated research?
- How does automation obscure failure modes in ways that make detection harder?
- Can automating failure absorption hide problems that governance needs to surface?
- How do traditional quality assurance methods fail for mutable AI outputs?
- Why is error rate alone misleading without strong contestability conditions?
- What happens when students encounter errors they cannot resolve through prompting alone?
- What specific failure modes must evaluation catch before deploying action-capable systems?
- Why do quiet failures reach deployment scale more often than loud ones?
- Which AI safety problems lack the scalar metrics autoresearch requires?
- Why do bug fixes carry more weight than hyperparameter tuning in pipelines?
- How does executable evaluation feedback sustain autonomous discovery at scale?
- What types of social situations cause all AI models to fail in identical ways?
- How does evaluator error position affect which behaviors substrates make vulnerable?
- What happens when error accumulation and preference signal collapse occur together?
- How do past research mistakes prevent future pivot loops from repeating them?
- How do autonomous pipelines identify and fix silent bugs in data pipelines?
- How do missing ground-truth exploits make it harder to identify genuine failures?
- What makes intermediate primitives matter more than final code execution success?
- What makes some frictions negligible while others block entire pathways?
- Why do some students restart entire projects instead of debugging incrementally?
- How do default fallback scores mask failures in evaluation harnesses?
- Why do frontier model failures in document editing go undetected by users?
- Are hedging markers in incorrect traces indicators of failed backtracking?
- Can reliable failure detection prevent optimization pressure against detectors?
- What breaks when a mis-synthesized verifier runs with high confidence?
- What does recovery look like as a formal part of AI design?
- What distinguishes an error bound from a forecast of system behavior?
- How does error avalanching differ from entropy collapse as a failure mode?
- What makes code inspectable feedback more reliable than natural language verification?
- What role does rigid output format play in function calling failure modes?
- What makes diverse failure modes more informative than single failure examples?
- When is GPT model interpretation most likely to diverge from user intent?
- Why do unresolved items cluster in structured patterns rather than randomly?
- How should tutor safety violations be ordered from gross to subtle?
- Why are closed AI systems harder to hold accountable than open ones?
- What distinguishes a component failure from a monitoring coverage failure?
- Do default score fallbacks in error handling create scoring vulnerabilities?
- When should a pipeline substitute defaults versus rejecting malformed outputs?
- What three independent failure points bottleneck traditional function calling systems?
- What role does opacity play in preventing effective challenge to system outputs?
- What is the generation-verification gap that predicts this failure mode?
- What baseline evidence distinguishes amplification from unchanged failure rates?
- What failure modes does the negative-space checklist generation method actually catch?
- Where do collider-type reasoning errors appear in real-world decisions?
- What status categories best represent user goal progress without penalizing external failures?
- What does tight coupling mean in normal accident theory for AI?