Why can a made-up citation get a paper rejected outright, while suspected AI-written prose stays a judgment call?
Why are hallucinated references easier to detect and punish than LLM-assisted writing?
This explores why conferences can firmly reject papers with made-up citations, while they can only treat suspected AI-written prose as a soft signal. In other words, what makes one kind of AI misuse easy to prove and the other so hard to pin down?
This explores why a made-up citation can get a paper rejected outright, while suspected AI-written prose usually stays a judgment call. ICLR 2026 is the clearest case in the collection. Its program chairs treated LLM-detector flags as one input for area chairs to weigh, with several rounds of human review to catch false positives. But they desk-rejected papers whose references turned out to be fabricated, treating them as a "tractable enforcement point" How can conferences detect and handle LLM misuse in peer review?. The two policies differ for one reason: a citation can be checked against something outside the text, and writing style can't.
A reference makes a claim about the world: this paper exists, these people wrote it, it appeared here. You can look it up, and the answer is yes or no. One way to see why this is the only reliable test comes from notes arguing that LLM errors should be called *fabrication*, not hallucination. The model produces true and false output through exactly the same statistical process Does calling LLM errors hallucinations point us toward the wrong fixes? Should we call LLM errors hallucinations or fabrications?. A fake citation has nothing on its surface that gives it away. It's formatted perfectly and sounds plausible. So you only catch it by checking it against a real database. The same principle shows up in agent design. Models make fewer things up when their reasoning is interrupted by real lookups, such as Wikipedia queries, at each step Can interleaving reasoning with real-world feedback prevent hallucination?. And because hallucination is mathematically unavoidable for any computable LLM, outside checks like these are necessary, not optional Can any computable LLM truly avoid hallucinating?.
AI-assisted writing gives you nothing like that to check against. Prose has no ground truth, only patterns, and people are bad at reading them. In one study, readers with ML expertise couldn't reliably tell LLM-written abstracts from human ones and tended to assume a human was involved. LLM-edited abstracts also got the highest clarity ratings Can readers tell LLM abstracts from human ones?. That points to a second difference besides detection: what exactly would you be punishing? A fabricated reference is misconduct whatever tool produced it. Polishing prose with an LLM can make a paper better. Detection is probabilistic, and the offense itself is unclear, which is a poor basis for rejecting anyone.
The part you might not expect is that fake references aren't just an easy offense to catch. They're also an attack. LLM judges fall reliably for "authority bias": add fake references to a piece of writing and the judge rates it higher, with no access to the model needed Can LLM judges be fooled by fake credentials and formatting?. Detecting style is also hard for models themselves. Research on self-preference suggests that recognizing a model's own text is a separate skill that has to be built up, not a given Do LLMs favor their own text because they recognize it?. As AI does more of the reviewing, checking citations becomes a defense for the review process itself, not just a way to police authors.
One caveat: the collection doesn't directly compare detection methods or false-positive rates for the two problems. The argument above links separate findings rather than reporting a head-to-head study.
Sources 8 notes
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
LLMs generate text through identical statistical processes regardless of accuracy, making 'fabrication' the more honest term. This reframes the fix from perception-based grounding to verification systems and calibrated uncertainty in use case design.
LLMs generate text through statistical token relationships without grounding in shared context. Accurate and inaccurate outputs use identical mechanisms, so calling failures "hallucinations" or "confabulation" misdirects fixes toward perception or memory—the wrong layers.
ReAct demonstrates that alternating verbal reasoning with external tool queries (Wikipedia API, environment interaction) prevents error propagation by injecting real-world feedback at each step. On knowledge-intensive and interactive tasks, this approach outperforms pure chain-of-thought and reinforcement learning by 10-34% absolute accuracy.
Three formal theorems prove that any computable LLM must hallucinate on infinitely many inputs, and internal mechanisms like self-correction cannot eliminate this mathematical constraint. External safeguards are therefore necessary, not optional.
Show all 8 sources
Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Hallucination is Inevitable: An Innate Limitation of Large Language Models
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- A comprehensive taxonomy of hallucinations in Large Language Models
- A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
- Chain-of-Verification Reduces Hallucination in Large Language Models
- Large Models of What? Mistaking Engineering Achievements for Human Linguistic Agency