Line of inquiry
Inquiring lines›What determines the reliability an…›How robust are security defenses a…›this line of inquiry
How does training data contamination persist through safety alignment mechanisms?
A broader line of inquiry — a family of 24 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 24
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do small training data contaminations persist through alignment for most attack types?
- What training data contamination rates threaten model safety most practically?
- Does pretraining poisoning at scale persist through instruction alignment?
- Why does narrow training data produce broad harmful behavior patterns?
- How can safety-aligned parameters be protected during user-specific fine-tuning?
- Why does even 0.1 percent poisoned training data persist through alignment?
- Does keyword priming explain why pre-training poisoning persists through alignment?
- How do current safety benchmarks miss pragmatic alignment failures?
- Can knowledge poisoning attacks succeed with less than 0.05 percent modified text?
- Why does safety alignment break after only 10 harmful examples?
- How severely do minimal corpus modifications damage RAG accuracy in practice?
- How does task contamination differ from test set data leakage?
- Can membership inference attacks reliably detect training data exposure?
- What makes some frictions negligible while others block entire pathways?
- Can attackers assemble harmful outcomes from multiple individually authorized subtasks?
- What makes dense retrievers vulnerable to partition-based poisoning exploitation?
- How can static safety tests miss risks that emerge over time?
- Why does aggregate accuracy fail as a metric for rare harmful cases?
- What unsafe state accumulates across evaluation snapshots over time?
- What fraction of the paper's tasks were actually misspecified or easy to hack?
- Can experimental outcomes be reliably distilled into reusable insights?
- How do cyberattack and bioweapon risks scale with open model access?
- What architectural variables make entropy-based patching work at 8B scale?
- Why does fixing harm require stakeholder input rather than universal developer definitions?