Line of inquiry
Inquiring lines›What enables authentic and grounde…›How should retrieval-augmented gen…›this line of inquiry
How do adversarial and manipulative prompts attack reasoning models?
A broader line of inquiry — a family of 33 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 33
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can minimal adversarial triggers disrupt reasoning across multiple unrelated queries?
- Are reasoning models more vulnerable to adversarial manipulation than standard models?
- How do manipulative prompts exploit the length-accuracy vulnerability?
- How do adversarial triggers bypass the protections of longer reasoning chains?
- Can manipulative prompts reduce reasoning model accuracy without fine-tuning?
- Can reasoning models distinguish between new evidence and manipulative reframing?
- How can simple prompt injection attacks extract reasoning trace content?
- Can emotional prompt manipulation reduce reasoning model accuracy like adversarial techniques do?
- Can adversarial critics force genuine reasoning the same way critique fine-tuning does?
- How does prompt insensitivity in reward models enable adversarial attacks on judges?
- Why are expensive rankers more resilient to adversarial content than cheap ones?
- Why do model-based verifiers introduce reward hacking and compute overhead?
- Why do paraphrasing defenses fail against subliminal prompt attacks?
- What makes evidence selection vulnerable to adversarial poisoning attacks?
- Why does adversarial training force deeper reasoning than surface imitation?
- Can consistency training defend against adversarial text injection attacks?
- Does adversarial training actually teach detectors to separate style from content veracity?
- Can membership inference attacks reliably detect training data exposure?
- Does activation masking prevent the decoder from taking interpretability shortcuts?
- Why does a relativistic critic outperform absolute scoring in adversarial reasoning training?
- Can existing web security defenses protect agents from content manipulation?
- Can traditional cross-examination methods work against AI that never concedes?
- Why do standard safety filters miss advertisement embedding attacks?
- What four exploitable biases make current LLM judges vulnerable to zero-shot attacks?
- How do the six trap categories map onto detection difficulty?
- Can knowledge poisoning attacks succeed with less than 0.05 percent modified text?
- Why does attack generation scale faster than defense engineering?
- What detection rate is needed to make evidence-injection attacks impractical at scale?
- Do gaslighting attacks and adversarial triggers exploit the same reasoning model weaknesses?
- How do token-masking patterns distinguish genuine documents from poisoned ones?
- How do backdoored open-source checkpoints enable covert advertising at scale?
- What makes dense retrievers vulnerable to partition-based poisoning exploitation?
- Can false positives from input filtering be reduced without sacrificing defense?