Can an AI learn to spot its own writing, and does getting better at that make it favor its own work?
Can models be trained to recognize their own generated text reliably?
This explores whether a language model can learn to tell its own writing apart from other text (human or other models), and what that ability does to how the model judges its own work.
This explores whether a language model can learn to tell its own writing apart from other text, and what happens once it can. The most direct answer in the collection is yes, at least for one kind of text: researchers fine-tuned models to recognize their own summaries, and it worked. The model's recognition got better. But the same experiment shows a cost. As recognition improved, the model also came to prefer its own summaries more, and the two rose together in a roughly straight line Do LLMs favor their own text because they recognize it?. The authors call this early causal evidence, not proof. If it holds up, a model that knows its own handwriting will favor its own handwriting.
The more surprising finding is that models may already recognize their own text without any training for it, and without saying so. Post-trained models are three to four times less uncertain when they read text they generated than when they read someone else's. That drop traces back to an internal signal about how surprising the input is Why do models produce less uncertain outputs on their own text?. Your own words are never surprising to you. So this kind of self-recognition shows up as confidence, not as an explicit judgment like "I wrote this." That helps explain a well-documented problem: models over-trust answers they produced themselves, because high-probability text feels correct when they evaluate it. The fix that works is to make the model compare its answer against a wider set of alternatives, so it isn't only agreeing with itself Why do models trust their own generated answers?.
This creates a tension for anyone using LLMs as judges or self-checkers. The model detects its own text best by feeling familiar with it, and that familiarity is exactly what distorts its judgment. LLM judges already fall for surface cues like fake credentials and rich formatting Can LLM judges be fooled by fake credentials and formatting?. Self-familiarity is a quieter version of the same weakness. A related line of work goes the other way: instead of asking a model to judge outputs after the fact, it trains self-evaluation into the model during training Can models learn to evaluate their own work during training?. In scientific discovery, other researchers simply give up on self-assessment and attach an external statistical model to estimate how good the LLM's candidates are Can language models reliably judge their own candidate quality?.
There is also a human side. AI text differs from human writing in ways statistics can measure, such as lexical diversity. Even trained linguists can't reliably spot it, and newer models drift further from human writing while getting harder to detect Can humans detect AI text if machines can measure it?. So the signal exists, but it sits below what readers notice. That gap matters because writers edit AI drafts only about 23% of the time, and lightly when they do Do writers actually edit AI-generated text before publishing?. Most AI text reaches readers close to unchanged.
One gap to be honest about: this collection doesn't have work on training models as reliable detectors of their own text in general, for example across topics, after human edits, or against other models' writing. What it shows is narrower and maybe more useful. Self-recognition can be trained and already shows up implicitly, but it is tied up with self-preference. A model that is better at recognizing its own writing may be a worse judge of it.
Sources 8 notes
Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.
Post-trained models produce 3-4x lower output entropy on their own generations, driven by an internal representation of input surprise that causally modulates confidence. This implicit self-recognition signal appears without being verbalized, encoded directly in the output distribution.
LLMs exhibit structural bias toward validating their own outputs because high-probability generated answers feel more correct during evaluation. Comparing answers against broader alternatives breaks this self-agreement loop.
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
Post-Completion Learning exploits unused sequence space after model output to train self-assessment capabilities during training while maintaining zero inference cost. The model learns to compute its own reward functions, internalizing evaluation rather than relying on external reward models.
Show all 8 sources
LLMs excel at generating valid candidates in structured spaces but cannot reliably assess their true value or uncertainty. Coupling them with Gaussian process surrogates fitted to real experimental data creates uncertainty-aware guidance for discovery.
LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.
Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- LLM Evaluators Recognize and Favor Their Own Generations
- Large Language Models Cannot Self-Correct Reasoning Yet
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Self-reflective Uncertainties: Do LLMs Know Their Internal Answer Distribution?
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- Think Twice Before Trusting: Self-Detection for Large Language Models through Comprehensive Answer Reflection
- Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries