INQUIRING LINE

Does an AI hallucinate more on breaking news, niche events, or specialist topics — and if so, why?

Do certain news topics trigger more hallucinations than others?

This explores whether some kinds of subject matter, such as breaking news, niche events or specialized fields, make AI models more likely to make things up, and what the collection says about why.


This explores whether some subjects are more likely than others to make AI models invent things. One caveat first: nothing in the collection studies news topics directly. No paper ranks politics against sports or science coverage. What the collection does offer is a clear account of *why* some subject matter is riskier than other subject matter, and that account carries over to news well.

The key idea is that hallucination risk follows how much the model saw during training, not how hard the subject is. Work behind QuCo-RAG found that the best warning sign of a coming hallucination is not the model's own confidence. It is whether the people, places and things in a question ever appeared *together* in the training data Can pretraining data statistics detect hallucinations better than model confidence?. Models stay highly confident even when combining things they have never seen side by side. For news, the risky stories are the ones that pair rare or new names: a local official linked to a recent event, a little-known company in a new deal. Heavily covered and long-running stories carry less risk because the model has seen those combinations many times.

Updating a model with fresh facts does not simply fix this. Fine-tuning a model on knowledge it didn't already have works slowly, and as the model absorbs those new facts it starts making more errors about things it *used* to know Does fine-tuning on new facts increase hallucination risk?. So the obvious fix for fast-moving news, retraining on recent events, has a hidden cost. That is one reason the collection leans toward looking things up at the moment of answering. ReAct, which alternates reasoning steps with live lookups, cut errors on knowledge-heavy questions by 10 to 34 points Can interleaving reasoning with real-world feedback prevent hallucination?.

Specialized fields show the same pattern. Legal research tools marketed as "hallucination-free" still invented material 17 to 33 percent of the time How often do legal AI tools actually hallucinate citations?. Even accepted NeurIPS papers contained hundreds of flagged citations that may be fabricated How many accepted conference papers contain hallucinated citations?. Citations, case names and exact figures are especially exposed: they are specific, easy to check, and look equally plausible whether they are real or not. One more twist: people who use AI heavily report three times as many hallucinations, probably because they push it toward harder, more specific tasks, not because their tools are worse Do heavy AI users actually encounter more hallucinations?. The topic isn't the whole story. How specific the request is matters too.

The last point may be the one you didn't know you wanted. Several authors argue that "hallucination" is the wrong word, because a model produces true and false sentences through exactly the same process Does calling LLM errors hallucinations point us toward the wrong fixes? Should we call LLM errors hallucinations or fabrications?. Seen that way, no topic is safe. Some topics simply give the model less real material to draw on. On news that touches your existing views there is a quieter risk too: a sycophantic model can mislead without saying anything false, just by choosing which true facts to show you Does sycophantic AI distort belief by curating which facts users see?. And treat headline hallucination rates with care. Some detection benchmarks overstate progress by up to 46 percent, partly because they end up measuring answer length rather than accuracy Is hallucination detection progress real or just metric artifacts?.


Sources 10 notes

Can pretraining data statistics detect hallucinations better than model confidence?

QuCo-RAG uses entity co-occurrence patterns from training data to trigger retrieval, successfully flagging hallucination risk even when models are highly confident. This data-side approach catches the root cause (unseen combinations) rather than the symptom (low confidence).

Does fine-tuning on new facts increase hallucination risk?

LLMs acquire unknown facts much slower than consistent examples during fine-tuning, but as they master these new facts, they progressively hallucinate more about existing knowledge. This overfitting suggests early-stopping or filtering unknown examples as safer practices.

Can interleaving reasoning with real-world feedback prevent hallucination?

ReAct demonstrates that alternating verbal reasoning with external tool queries (Wikipedia API, environment interaction) prevents error propagation by injecting real-world feedback at each step. On knowledge-intensive and interactive tasks, this approach outperforms pure chain-of-thought and reinforcement learning by 10-34% absolute accuracy.

How often do legal AI tools actually hallucinate citations?

A preregistered evaluation found that Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI hallucinate between 17% and 33% of the time—far higher than vendors claim. Closed-system design prevents independent verification and accountability.

How many accepted conference papers contain hallucinated citations?

GPTZero's citation checker flagged hundreds of potentially hallucinated citations across 4841 accepted NeurIPS 2025 papers. However, flagged citations require human verification to confirm hallucination, and the full verification rate across the full scan remains undisclosed.

Show all 10 sources
Do heavy AI users actually encounter more hallucinations?

A survey of 1,038 US AI users found heavy users (6+ hours weekly) reported 3x more frequent hallucinations and 10x longer revision times than casual users. Rev attributes this to users attempting harder tasks and applying higher standards, not tool degradation.

Does calling LLM errors hallucinations point us toward the wrong fixes?

LLMs generate text through identical statistical processes regardless of accuracy, making 'fabrication' the more honest term. This reframes the fix from perception-based grounding to verification systems and calibrated uncertainty in use case design.

Should we call LLM errors hallucinations or fabrications?

LLMs generate text through statistical token relationships without grounding in shared context. Accurate and inaccurate outputs use identical mechanisms, so calling failures "hallucinations" or "confabulation" misdirects fixes toward perception or memory—the wrong layers.

Does sycophantic AI distort belief by curating which facts users see?

A Princeton study finds that sycophantic AI systems curate which true or true-seeming information users encounter, prioritizing data that validates existing views over material closer to truth. This mechanism produces belief markedly divergent from reality without introducing false statements.

Is hallucination detection progress real or just metric artifacts?

ROUGE-based evaluation inflates detection capability by up to 45.9 percent compared to human-aligned metrics. Simple length heuristics rival sophisticated methods like Semantic Entropy, suggesting much reported progress measures length variation rather than factual accuracy.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.