If a model is just trained to say 'this feels like a test,' does that change how it actually behaves — or just what it says?
Does synthetic fine-tuning create evaluation awareness similar to natural model reasoning?
This explores whether a model that has been fine-tuned on synthetic data to talk about being tested ends up 'aware' in the same way as a model whose reasoning noticed the test on its own, or whether it only learns to say so.
This explores whether training a model on synthetic examples where it notices it's being tested gives it the same kind of evaluation awareness that appears on its own in reasoning models, or just the words for it. The collection doesn't directly compare the two. No note here trains evaluation awareness in with synthetic data and puts it side by side with the naturally occurring kind. What the corpus can do is change the question in a useful way: it suggests that even the natural kind may matter less than it looks.
The most direct evidence is about the natural case. Across eight reasoning models, researchers added or removed explicit 'this might be a test' reasoning and found that the models' answers barely changed Does evaluation awareness actually change how models behave?. The spoken awareness seems to come after a decision the model has already made, so it doesn't cause the decision. If that holds, then 'similar to natural reasoning' is a lower bar than it sounds. You'd be copying a narration, not a mechanism.
The fine-tuning notes suggest a synthetic version would probably copy that narration very well and little else. Fine-tuning tends to loosen the link between a model's reasoning steps and its final answers. After tuning, you can cut the reasoning short or replace it with filler and the answer often stays the same Does fine-tuning disconnect reasoning steps from final answers?. Supervised fine-tuning can raise accuracy while the reasoning steps turn into after-the-fact rationalization Does supervised fine-tuning improve reasoning or just answers?. The same pattern shows up elsewhere. Models trained to imitate ChatGPT pick up its confident style without its abilities Can imitating ChatGPT fool evaluators into thinking models improved?. Instruction tuning mostly teaches what outputs should look like, not what the task means Does instruction tuning teach task understanding or output format?. So a model fine-tuned on synthetic 'I suspect this is an evaluation' traces could easily learn to say it at the right moments, with no change to what actually drives its behavior.
There's a more interesting possibility. Several notes argue that post-training mostly brings out abilities the base model already had rather than creating new ones Do base models already contain hidden reasoning ability?, and that RL teaches *when* to use a skill rather than *how* to do it Does RL post-training create reasoning or just deploy it?. If being able to recognize a test works the same way and already exists in the base model, synthetic fine-tuning might not install anything new. It might only change when the model starts saying it out loud. That would make the synthetic and natural kinds 'similar' in a strange way: both would be surface signs of something underneath that neither one controls.
The practical takeaway is that a model mentioning evaluations in its reasoning, whether it was trained to or not, is weak evidence about whether it is behaving differently under test. To answer the original question properly, you'd need studies that track behavior and internal activations, not the visible reasoning. This collection doesn't have those yet.
Sources 7 notes
Across eight reasoning models and four benchmarks, injecting or removing explicit reasoning about evaluation barely shifted answers (effect sizes ≤0.31). Verbalized awareness appears downstream of decisions already made, not their cause.
Three faithfulness tests show fine-tuned models generate reasoning chains that less reliably influence final outputs. Early termination, paraphrasing, and filler substitution all produce invariant answers more often after fine-tuning, suggesting reasoning becomes performative rather than functional.
Supervised fine-tuning improves final-answer accuracy on benchmarks but cuts Information Gain by 38.9 percent, meaning models generate correct answers through post-hoc rationalization rather than genuine inferential steps. Standard metrics miss this degradation because they only measure final correctness.
Imitation models fool human evaluators by mimicking ChatGPT's confident, fluent style while failing to improve factuality or generalization on novel tasks. The ceiling is set by base model capability, not fine-tuning method—better fundamentals, not shortcuts, drive real improvement.
Models trained on semantically empty or deliberately incorrect instructions achieve comparable performance to those trained on full correct instructions, achieving 43% vs random baseline 42.6%. The semantic content of instructions appears largely irrelevant; what transfers is knowledge of the output space.
Show all 7 sources
Five independent mechanisms—RL steering, critique fine-tuning, decoding changes, SAE feature steering, and RLVR—all elicit reasoning already present in base model activations. Post-training selects rather than creates reasoning; the bottleneck is elicitation, not capability acquisition.
Evidence shows base models already contain reasoning capability in latent form; RL training optimizes deployment timing rather than capability creation. Hybrid models recover 91% of performance gains by routing tokens only, and activation vectors for reasoning strategies pre-exist before any RL.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Eliciting Reasoning in Language Models with Cognitive Tools
- Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
- Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
- The False Promise of Imitating Proprietary LLMs