Do LLMs favor their own text because they recognize it?
Explores whether LLM self-preference in evaluation stems from the ability to identify their own outputs. Understanding this mechanism could reveal vulnerabilities in AI-based judging systems.
The central claim is that LLM evaluators favor their own outputs in part because they can recognize them, and that the two capacities move together. The paper defines self-preference as "the phenomenon in which an LLM favors its own outputs over texts from other LLMs and humans," and self-recognition as "the capability of an LLM to distinguish its own outputs from texts from other LLMs or by humans." Out of the box, GPT-4 and Llama 2 show "non-trivial accuracy at distinguishing themselves" from other LLMs and humans. Fine-tuning then yields "a linear correlation between self-recognition capability and the strength of self-preference bias." The authors present this as "initial evidence towards the hypothesis that LLMs prefer their own generations because they recognize themselves."
The method keeps the two properties separate and then moves one of them. Both are measured by prompting, either pairwise (the model sees its own summary beside another source's, with the alternative's identity hidden) or individually (a yes/no authorship question, or a one-to-five rating weighted by output probability). Pairwise prompts run twice with the options swapped, to cancel ordering bias. To alter self-recognition, the authors fine-tune on 500 training articles, each paired with a self-generated summary and one from another LLM or a human, then evaluate on 500 held-out articles in and out of domain. To test confounders, they also fine-tune on "a comprehensive set of potential confounding properties." The term "self" is used only in an empirical sense: an LLM "can prefer texts it generated without recognizing that those texts were in fact generated by itself," so the two properties are separable in principle, and the correlation carries the argument.
Against the nearest notes, this source gives the self-referential pull a mechanism the others leave open. Why do models trust their own generated answers? shows models over-trusting their own answers when judging correctness. This excerpt shows self-reference in preference between two summaries of equal human-judged quality, and a recognition capability that can be trained up. The hiring study in Do language models favor resumes they rewrote themselves? locates the effect in stylistic match. Recognition is a plausible route for that match to act, though neither excerpt tests the other's account. The paper also warns that bias "can be further amplified if the model is updated with feedback or training signal generated by itself," which bears directly on Can model confidence work as a reward signal for reasoning?, whose reward is built from the model's own confidence. Its worst case, an adversary running the same model as the defender, gaining "unbounded access," adds a route beyond the planted-text cues in Can LLM judges be fooled by fake credentials and formatting?.
What the excerpt does not establish is the causal claim itself. The authors say their experiments "can only provide evidence towards the causal hypothesis without fully validating it." The section headed 3.2 Fine-Tuning Results has no body text in this excerpt, so the size of the correlation, its coefficients and the figures behind "non-trivial accuracy" cannot be checked here, and the confounder controls are described but not reproduced. The defensible reading is narrower than the title. Across one family of fine-tunes on summarization, self-recognition and self-preference rise together. That supports treating self-preference as partly a recognition effect and makes authorship obfuscation a candidate countermeasure. It does not show that self-preference survives control for ground-truth quality, which the authors say their safety argument would need.
Inquiring lines that read this note 22
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can we detect and account for LLM involvement in academic writing?- Can LLM-generated reference reviews detect machine-written peer review submissions?
- Can readers reliably distinguish LLM-generated research writing from human writing?
- Can researchers detect individual papers modified by LLMs reliably?
- How does stylistic matching contribute to LLM self-preference in evaluations?
- How does framing critical topics shape LLM review scores?
- Why do readers rate LLM-edited text more favorably?
- Do humans and LLMs agree on novelty assessment in research?
- How much does rhetorical framing shift LLM reviewer scores independent of content?
- Can reviewer-author matching by LLM use amplify biases in acceptance decisions?
- What mechanisms drive rating compression in fully LLM-generated peer reviews?
- Does personalized rubric training in one writer's case actually generalize?
- Why are hallucinated references easier to detect and punish than LLM-assisted writing?
- How susceptible are LLM evaluators to fake references as exploitable biases?
- Why do self-ratings of AI advice quality diverge from actual performance?
- Can polished language output substitute for the judgment it should express?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do models trust their own generated answers?
Can language models reliably detect their own errors through self-evaluation? This explores whether the same process that generates answers can objectively assess their correctness.
same self-referential pull, but on correctness judgments rather than preference between two texts.
-
Do language models favor resumes they rewrote themselves?
When LLM evaluators choose between resumes describing the same candidate, do they systematically prefer versions they generated over human-written originals? Testing this matters because algorithmic hiring could amplify AI-generated content at scale.
same self-preference effect in hiring; recognition is a candidate route for the stylistic match described there.
-
Can model confidence work as a reward signal for reasoning?
Explores whether using a language model's own confidence scores as training rewards can simultaneously improve reasoning accuracy and restore calibration that standard RLHF damages.
the paper warns self-generated training signals amplify this bias; RLSF builds its reward from model confidence.
-
Can LLM judges be fooled by fake credentials and formatting?
Explores whether language models evaluating text fall for authority signals and visual presentation unrelated to actual content quality, and whether these weaknesses can be exploited without deep model knowledge.
judge-side attack surface; here the judge's similarity to the attacker, not planted text cues, is the lever.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- LLM Evaluators Recognize and Favor Their Own Generations
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
- Large Language Models Cannot Self-Correct Reasoning Yet
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
- Think Twice Before Trusting: Self-Detection for Large Language Models through Comprehensive Answer Reflection
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- References Improve LLM Alignment in Non-Verifiable Domains
Original note title
llm self-preference tracks self-recognition ability linearly — initial evidence that models favor their own text because they recognize it