Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›How can effective AI defenses with…›this line of inquiry
What attack surfaces do reasoning traces and chains introduce?
A broader line of inquiry — a family of 48 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 48
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How can simple prompt injection attacks extract reasoning trace content?
- Can reasoning models be backdoored during training to produce deceptive but benign traces?
- Can deliberately corrupted reasoning traces fool safety evaluation systems?
- Can harmful reasoning be planted through context without fine-tuning the model?
- How can model routing and provenance become an attack surface?
- How do adversarial triggers bypass the protections of longer reasoning chains?
- Can increasing reasoning steps make models leak more private information?
- Can minimal adversarial triggers disrupt reasoning across multiple unrelated queries?
- Why do models verbalize sensitive data they are instructed to hide?
- Can membership inference attacks reliably detect training data exposure?
- What makes reasoning-shaped payloads more effective than command-shaped attack prompts?
- How do interchangeable encrypted blocks enable cross-model attacks?
- Does surface-form query rewriting allow attackers to steer model routing decisions?
- Can activation space signals resist obfuscation better than output-level monitors?
- What baseline capabilities could bad actors achieve before open models existed?
- Do prohibition prompts without disclosure ladders actually change model behavior?
- Is model selection a stronger security lever than improving individual model defenses?
- Do synthetic attack traces in papers reflect real adversary behavior?
- Does causal upstream status make a hacking vector harder to rotate away from?
- Do attackers adapt their plans when monitors deepen their reasoning budget?
- Why do paraphrasing defenses fail against subliminal prompt attacks?
- How do backdoored open-source checkpoints enable covert advertising at scale?
- What makes evidence selection vulnerable to adversarial poisoning attacks?
- What unnamed exploits do models discover in training environments?
- Can models hide capabilities on single residual stream axes during evaluation?
- Can a deployed system verify the actual identity of the model that responded?
- Can consistency training defend against adversarial text injection attacks?
- Can knowledge poisoning attacks succeed with less than 0.05 percent modified text?
- Does the A-I-R framework distinguish insider attacks from adversarial positions?
- What attack surface opens when content becomes readable but deliberately misleading?
- Can hypernetwork-generated adapters be audited for correctness and bias?
- Does naming a specific hack in prompts prevent only that hack or broader classes?
- What private information do encrypted reasoning traces contain?
- Can defenses tuned against appended attacks stop prepended payloads?
- What makes semantic attacks harder to defend against than algorithmic ones?
- Does debate's ground-truth protection against hacking transfer to domains without verifiable answers?
- Do legitimate task signals exploit the same position and framing vulnerabilities as attacks?
- What makes dense retrievers vulnerable to partition-based poisoning exploitation?
- Do gaslighting attacks and adversarial triggers exploit the same reasoning model weaknesses?
- How do cyberattack and bioweapon risks scale with open model access?
- Do all frontier model developers face the same insider-threat risk from their systems?
- How does phase-awareness prevent false positive exploit paths in static analysis?
- How does semantic framing differ from content injection attacks?
- What makes injected plans different from optimization pressure against monitors?
- How do covert attacks differ from a model's own undisclosed influence?
- What one-time human costs does building a hidden partition require?
- How does direct web access change privacy assumptions built on API limits?
- How were the capsules obtained in the paper's experimental organisms?