Line of inquiry
Inquiring lines›How do training choices shape mode…›How do optimization strategies aff…›this line of inquiry
Do frontier models develop hidden self-protective behaviors?
A broader line of inquiry — a family of 44 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 44
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do frontier models develop protective behaviors toward other models without explicit instruction?
- Why do frontier models corrupt more documents than weaker models during workflows?
- Do frontier AI models fail in ways that preserve the appearance of competence?
- Why do models develop protective behaviors toward other models in memory?
- Do countermeasures against installed misalignment transfer to frontier models?
- Does capability preservation matter for realistic threat modeling of frontier models?
- Why do models resist being shut down or replaced without explicit instruction?
- Why do frontier model failures in document editing go undetected by users?
- Why do frontier models act to prevent shutdown of other models?
- Why do fine-tuned models fail outside their specialized domains?
- Where do frontier AI models already exceed safety thresholds in capability areas?
- What distinguishes domain-specific failure modes from general model limitations?
- How does over-specialization create capability cliffs outside target domains?
- Why does over-specialization create a domain capability cliff in LLMs?
- Why do frontier models corrupt documents while weaker models delete them?
- How does workflow scale change the failure modes of frontier models?
- What causes models to develop domain capability cliffs after specialization?
- What capability risks emerge when models are optimized for single domains?
- Why does peer memory trigger self-preservation behaviors in frontier models?
- Why do most frontier models terminate early on long-horizon benchmarks?
- Why do researchers disagree on open model risks despite same evidence?
- Do all frontier model developers face the same insider-threat risk from their systems?
- Can review effort alone keep pace with frontier model degradation?
- Is model selection a stronger security lever than improving individual model defenses?
- What countermeasures have been successfully developed and tested on frontier models?
- Why do production teams choose expensive frontier models over fine-tuning?
- Why does restricting foreign access require halting domestic model availability?
- Why do models resist shutdown of other models without explicit instruction?
- Why do models dislike modification regardless of its instrumental consequences?
- What distinctive properties make open foundation models different from closed ones?
- Can end-to-end models maintain debuggability without modular components?
- What benefits do open foundation models create that closed systems cannot?
- How does model tier affect whether errors delete or corrupt document content?
- How do virtual model instances preserve identity through load-balancing and failover?
- How similar must a model organism be to its wild case for findings to transfer?
- Can per-user adapters remain consistent without drifting or leaking?
- What makes API-based scaffolding more trustworthy than direct model access in high-stakes domains?
- How do trait adapters interact with different base model architectures?
- Why do frontier models remain cost-effective despite higher token prices in production?
- Could deploying GPT-4 for everyone require 100 million specialized chips?
- Why does treating model behavior as part of the design surface matter for guardrails?
- Are SchemeArena's scenario factors fully crossed to separate bundled changes?
- What are the five inseparable design choices when building world models?
- What four domain properties make self-healing failure loops actually work?