Do AI assistants reliably answer questions about news?
A major cross-country study tested whether AI assistants like ChatGPT and Gemini accurately handle news queries. Understanding this matters because many people may turn to these tools for current events.
The European Broadcasting Union, working with the BBC, coordinated what the report calls "one of the largest cross-market evaluations of its kind": 22 public service media organizations in 18 countries, working in 14 languages, had their professional journalists assess how ChatGPT, Copilot, Gemini, and Perplexity answer questions about news and current affairs. Journalists evaluated more than 3,000 AI responses against accuracy, sourcing, distinguishing opinion from fact, and providing context. The finding was stark: "almost half of all AI answers had at least one significant issue," a third showed "serious sourcing problems," and a fifth contained "major accuracy issues, such as hallucinated and/or outdated information." The report states the problem is not confined to any one assistant, language, or country — it found AI "routinely misrepresents news content, no matter which language, territory, or AI platform is tested."
The study's design is what gives the finding its force: it extends an earlier BBC study that had already flagged inaccuracies, and was built specifically to test whether that earlier problem was an isolated glitch or a systemic feature of how these assistants handle news queries. By recruiting journalists across 18 countries and 14 languages to apply the same evaluation criteria, the EBU could rule out the possibility that errors were an artifact of English-language training data, a single market's news ecosystem, or one vendor's particular failure mode. The uniformity of the result across markets and platforms is the evidence for systemic causation rather than isolated incidents.
This sits alongside Does AI fact-checking actually help people spot misinformation? as a second, independent line of evidence that AI-mediated engagement with news degrades reliability — but through a different mechanism. That study measured downstream belief effects when an AI fact-checker mislabels existing claims; this one measures upstream generation errors when an assistant is asked directly about current events, with sourcing and hallucination as the dominant failure modes rather than mislabeling. It also echoes, in a different domain, the finding in Do AI writing tools improve online discussion or degrade it? that AI involvement in informational exchange can erode quality even where it does not reduce use — here the erosion is in factual reliability of a news source rather than discourse authenticity.
The excerpt does not give the evaluation rubric's scoring thresholds, inter-rater agreement, or a breakdown of error rates by assistant, language, or topic, so it cannot show whether one platform is meaningfully better than the others or whether certain news topics (breaking news versus established fact, for instance) drive the sourcing and hallucination rates disproportionately. What it does establish, at the strength of a 3,000-response, 18-country sample evaluated by working journalists against a shared rubric, is that treating any current mainstream AI assistant as a reliable first stop for news questions carries a roughly coin-flip chance of a significant error — a finding strong enough to generalize across the four platforms tested, though not beyond them.
Inquiring lines that read this note 6
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can AI systems reliably guide voters without introducing political bias?- Does ChatGPT displace search engines or question-and-answer platforms?
- What accuracy do AI chatbots actually provide on election topics?
- Which AI news assistant performs better than the others in this study?
- How often do chatbot news users actually return to original reporting?
- Do other AI assistants perform similarly on voter election questions?
- What makes voting-advice tools like Kieskompas more reliable than chatbots?
Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does AI fact-checking actually help people spot misinformation?
An RCT tested whether AI fact-checks improve people's ability to judge headline accuracy. The results reveal asymmetric harms: AI errors push users in the wrong direction more than correct labels help them.
a second, independent mechanism by which AI-mediated news engagement degrades reliability
-
Do AI writing tools improve online discussion or degrade it?
When AI assists with comments and replies, does it benefit both people writing and reading? A controlled experiment tested whether AI tools enhance or harm the quality and authenticity of online conversations.
parallel finding that AI involvement in informational exchange can erode quality without reducing use
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- News Integrity in AI Assistants
- Artificial intelligence is ineffective and potentially harmful for fact checking
- AI and Elections: How Well Do AI Platforms Answer Voter Questions?
- Digital News Report 2026
- Emerging uses of AI chatbots for news and what it means for journalism (Digital News Report 2026)
- How AI Is Changing Search Behaviors
- AI Now Writes as Many Online Articles as Humans
- News Source Citing Patterns in AI Search Systems
Original note title
EBU and BBC's study of 3,000 AI responses finds AI assistants misrepresent news regardless of language, territory, or platform