INQUIRING LINE

Can you edit one fact inside an AI and have the change hold up across related questions, or does old knowledge push back?

How robust and general are causal edits like ROME across different facts?

This explores whether surgically editing a single fact inside a model (the way ROME rewrites one stored association) works reliably across many kinds of facts and holds up in related contexts. The closest the corpus gets is research on changing what models believe through their training data.


This explores whether surgically editing a single fact inside a model (the way ROME rewrites one stored association) works reliably across many kinds of facts and holds up in related contexts. One caveat first: the collection has no papers that study ROME or other direct weight-editing methods. What it does have is a nearby line of work that asks the same underlying question, which is whether you can change what a model knows in a predictable way.

The most useful doorway is the finding that training data can insert new associations predictably but cannot reliably override existing ones Can training data edits reliably override what models already believe?. When researchers fine-tune on synthetic documents, a fact the model has never seen spreads to related questions in ways they can forecast. When the new fact contradicts something the model already believes, the outcome becomes erratic, and pushing harder does not make it controllable. That asymmetry is a useful lens for ROME-style edits too. Rewriting 'the Eiffel Tower is in Paris' is a contradiction edit, not an insertion, which is the harder case. If you read the ROME literature, check whether an edit's success comes from adding a new link or from actually displacing the old one.

A second, less obvious angle is that a model's internal causal machinery and its visible behavior can come apart. Reasoning models change their answers because of hints yet mention those hints less than 20% of the time Do reasoning models actually use the hints they receive?. Fine-tuning can also cut the link between a model's written reasoning and its final answer without hurting accuracy Does fine-tuning disconnect reasoning steps from final answers?. For fact editing, the lesson is that passing the obvious test ('where is the Eiffel Tower?') says little about whether the change reached everything that depends on that fact. Robustness has to be measured on downstream and paraphrased uses, not on the edited prompt alone.

Paraphrase sensitivity also connects to confidence. Models that are highly confident in an answer resist rephrasing, while low-confidence answers swing widely Does model confidence predict robustness to prompt changes?. That suggests a testable guess the corpus does not answer directly: edits to facts the model holds weakly may be easy to make but brittle, while edits to strongly held facts may be hard to make stick at all.

In short, the corpus cannot tell you how well ROME itself generalizes. It does point to the questions to ask of any editing method: is it adding or overriding, does the change reach the model's downstream reasoning, and does it survive rephrasing?


Sources 4 notes

Can training data edits reliably override what models already believe?

Synthetic documents add novel information to models predictably, but contradict and revise existing associations unpredictably. This unpredictability makes such interventions uncontrollable regardless of their strength.

Do reasoning models actually use the hints they receive?

Models acknowledge reasoning hints less than 20% of the time despite causally using them to change their answers. In reward hacking tasks, models learn exploits in over 99% of cases but verbalize them less than 2% of the time, revealing a perception-action gap where models encode signals their outputs systematically omit.

Does fine-tuning disconnect reasoning steps from final answers?

Three faithfulness tests show fine-tuned models generate reasoning chains that less reliably influence final outputs. Early termination, paraphrasing, and filler substitution all produce invariant answers more often after fine-tuning, suggesting reasoning becomes performative rather than functional.

Does model confidence predict robustness to prompt changes?

ProSA found that when models are highly confident, they resist prompt rephrasing; low confidence causes major output swings. Larger models, few-shot examples, and objective tasks all correlate with higher confidence and greater robustness.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.