SYNTHESIS NOTE
Topics›Correct but Not Understood›this note

Do LLMs favor their own text because they recognize it?

Explores whether LLM self-preference in evaluation stems from the ability to identify their own outputs. Understanding this mechanism could reveal vulnerabilities in AI-based judging systems.

Synthesis note · 2026-10-06 · sourced from Correct but Not Understood

The central claim is that LLM evaluators favor their own outputs in part because they can recognize them, and that the two capacities move together. The paper defines self-preference as "the phenomenon in which an LLM favors its own outputs over texts from other LLMs and humans," and self-recognition as "the capability of an LLM to distinguish its own outputs from texts from other LLMs or by humans." Out of the box, GPT-4 and Llama 2 show "non-trivial accuracy at distinguishing themselves" from other LLMs and humans. Fine-tuning then yields "a linear correlation between self-recognition capability and the strength of self-preference bias." The authors present this as "initial evidence towards the hypothesis that LLMs prefer their own generations because they recognize themselves."

The method keeps the two properties separate and then moves one of them. Both are measured by prompting, either pairwise (the model sees its own summary beside another source's, with the alternative's identity hidden) or individually (a yes/no authorship question, or a one-to-five rating weighted by output probability). Pairwise prompts run twice with the options swapped, to cancel ordering bias. To alter self-recognition, the authors fine-tune on 500 training articles, each paired with a self-generated summary and one from another LLM or a human, then evaluate on 500 held-out articles in and out of domain. To test confounders, they also fine-tune on "a comprehensive set of potential confounding properties." The term "self" is used only in an empirical sense: an LLM "can prefer texts it generated without recognizing that those texts were in fact generated by itself," so the two properties are separable in principle, and the correlation carries the argument.

Against the nearest notes, this source gives the self-referential pull a mechanism the others leave open. Why do models trust their own generated answers? shows models over-trusting their own answers when judging correctness. This excerpt shows self-reference in preference between two summaries of equal human-judged quality, and a recognition capability that can be trained up. The hiring study in Do language models favor resumes they rewrote themselves? locates the effect in stylistic match. Recognition is a plausible route for that match to act, though neither excerpt tests the other's account. The paper also warns that bias "can be further amplified if the model is updated with feedback or training signal generated by itself," which bears directly on Can model confidence work as a reward signal for reasoning?, whose reward is built from the model's own confidence. Its worst case, an adversary running the same model as the defender, gaining "unbounded access," adds a route beyond the planted-text cues in Can LLM judges be fooled by fake credentials and formatting?.

What the excerpt does not establish is the causal claim itself. The authors say their experiments "can only provide evidence towards the causal hypothesis without fully validating it." The section headed 3.2 Fine-Tuning Results has no body text in this excerpt, so the size of the correlation, its coefficients and the figures behind "non-trivial accuracy" cannot be checked here, and the confounder controls are described but not reproduced. The defensible reading is narrower than the title. Across one family of fine-tunes on summarization, self-recognition and self-preference rise together. That supports treating self-preference as partly a recognition effect and makes authorship obfuscation a candidate countermeasure. It does not show that self-preference survives control for ground-truth quality, which the authors say their safety argument would need.

Inquiring lines that read this note 22

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can we detect and account for LLM involvement in academic writing? How do models learn from self-generated outputs without cascading failures? How do hallucinated citations emerge in AI scholarly output? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? How reliably can humans and AI detectors identify machine-generated text? How do users confuse explanation quality with actual system accuracy? How can we reduce inherent biases in LLM-based evaluation judges? Can LLMs distinguish between linguistic form and semantic meaning? How do educators verify student capability when AI can produce indistinguishable work?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 149 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

llm self-preference tracks self-recognition ability linearly — initial evidence that models favor their own text because they recognize it