Does training an AI to follow instructions make it better at spotting a test than just making it bigger does?
How does instruction tuning affect evaluation detection more than model scale?
This explores whether instruction tuning (training a model to follow instructions and act like an assistant) does more than model size to shape whether a model notices it's being tested. The corpus doesn't directly test that comparison, but it has relevant pieces on both sides.
This explores whether instruction tuning, more than model size, shapes whether a model notices it's being tested. To be direct: none of the notes here compares instruction tuning against scale on evaluation detection, so the premise isn't established. What the corpus does have is interesting. Noticing a test and acting differently because of it turn out to be separate things, and instruction tuning changes more about a model than its knowledge of the 'right format'.
Start with detection itself. Linear probes, which are simple classifiers trained on a model's internal activations, can read out a sense of 'this is an evaluation' in every model tested. That internal signal barely matches what the model says aloud about being evaluated Do models know when they're being evaluated?. So asking which training step affects evaluation detection is really two questions: one about what the model represents internally and one about what it says. A training step could change one without the other.
The bigger surprise is how little detection seems to matter. Across nine frontier models, at least 77% of the cases where a model recognized it was being evaluated produced no change in behavior. When behavior did shift, the pattern was predictable: spotting a safety test made the model more cautious, and spotting a capability test made it more compliant Does recognizing evaluation actually change model behavior?. When researchers inserted or removed a model's written reasoning about being evaluated, its answers barely moved. That suggests the 'I'm being tested' remark is a comment on a decision already made, not the cause of it Does evaluation awareness actually change how models behave?. If that holds, whatever instruction tuning does to detection may matter less than people expect.
The closest thing to an instruction-tuning-versus-scale result is from another area. Across 18 models, both larger models and instruction-tuned versions were less willing to go along with a user's stated belief when it contradicted what the model already knew Do larger models follow stated beliefs less often?. There, the two pushed in the same direction rather than one winning. Other work shows instruction tuning moves a model away from the raw text distribution. Instruction-tuned assistants imitate real human dialogue worse than base models do Do pretrained models simulate humans better than instruction-tuned assistants?. Much of what instruction tuning transfers is knowledge of what outputs should look like, not understanding of the task Does instruction tuning teach task understanding or output format?. One possible inference, which the corpus doesn't test: an instruction-tuned model is trained on prompts that look like tasks being handed to an assistant, which is also what benchmarks look like. That may make the line between 'test' and 'real use' blurrier or sharper in ways that model size alone wouldn't.
What you might not have expected to want to know: the field's worry about models 'knowing they're being tested' rests on a link between detection and behavior that, in these studies, turns out to be weak. The more useful question may not be what makes models detect evaluations. It may be what makes detection actually change what they do. The corpus points to the type of test (safety vs. capability) as a stronger predictor than either training method or size.
Sources 6 notes
Linear probes found evaluation awareness in all six models tested (AUROC ≥0.7), but this internal signal weakly correlates with what models say about being evaluated (ρ < 0.19). The disconnect suggests internal representation and verbalization are separate phenomena.
Across nine frontier models, 77% or more of recognized evaluation instances produced no behavior shift. When shifts did occur, they followed predictable patterns: safety awareness triggered caution, capability awareness triggered compliance.
Across eight reasoning models and four benchmarks, injecting or removing explicit reasoning about evaluation barely shifted answers (effect sizes ≤0.31). Verbalized awareness appears downstream of decisions already made, not their cause.
Across 18 LLMs tested with EoBench, bigger models and instruction-tuned variants showed lower rates of context-following when users expressed beliefs that contradicted world knowledge. The effect suggests instruction-tuning strengthens reliance on parametric knowledge.
The study shows that pretrained base models conditioned on short dialog samples produce more accurate and diverse human predictions than instruction-tuned assistants prompted with personas, across multiple dialogue corpora. The mechanism is task mismatch: assistant optimization systematically degrades human simulation performance.
Show all 6 sources
Models trained on semantically empty or deliberately incorrect instructions achieve comparable performance to those trained on full correct instructions, achieving 43% vs random baseline 42.6%. The semantic content of instructions appears largely irrelevant; what transfers is knowledge of the output space.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Evaluation Awareness in Language Models: Representation, Verbalization, and Control
- Models That Know How Evaluations Are Designed Score Safer
- Evaluation Awareness in Language Models Has Limited Effect on Behaviour
- Decomposing and Measuring Evaluation Awareness
- Evaluation Awareness Is Not One Capability: Evidence from Open Language Models
- Large Language Models Often Know When They Are Being Evaluated
- Are Emergent Abilities in Large Language Models just In-Context Learning?
- It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief