A huge BBC/EBU study graded 3,000 AI news answers — did any chatbot actually come out ahead, or do they all struggle equally?
Which AI news assistant performs better than the others in this study?
This explores whether the big EBU/BBC study of AI news answers found one assistant (ChatGPT, Copilot, Gemini, Perplexity) clearly more reliable than the rest, and what the corpus suggests about why the answer might matter less than it seems.
This explores whether one AI assistant came out clearly ahead in the large public-broadcaster study of AI news answers. The short answer from the collection: the study's main point is not that one assistant wins. It found that every major platform, in every language tested, gets the news wrong often. In Do AI assistants reliably answer questions about news?, journalists from 22 broadcasters in 18 countries graded more than 3,000 answers. Nearly half had a significant error, a third had serious sourcing problems, and a fifth had major accuracy failures such as made-up details. The note in this collection doesn't give a per-assistant ranking, so if you want to know which product scored best, the original report is the place to look. What the corpus does tell you is that this is a problem with the whole category, not with one bad product.
The sourcing finding helps explain why. Many of the failures weren't wrong facts so much as wrong or missing attribution: the assistant couldn't show where a claim came from, or pointed to a source that didn't say it. A related line of research, Do language models know what they don't know about users?, suggests a deeper cause. These systems have no explicit way to track what they don't know, so they fill gaps with confident text. When researchers gave models a structured list of known unknowns, hallucination dropped by about half. That points to a fix that has to do with how the model handles uncertainty, not with which brand you pick.
The more surprising part is how people behave anyway. Chatbot news use is rising, from 7% to 10% weekly across dozens of markets, even though only about 20% of people trust chatbot answers, compared with 37% for news in general (Why do AI chatbots gain news users but lose their trust?). Among people who already use chatbots for news, trust is much higher: 44% versus 17% for non-users (Does trust in AI chatbots drive news-seeking behavior?). So the people relying on these tools are the ones most inclined to believe them, which is exactly the group the error rates should worry. There is one counterweight. Panel data shows people usually open an assistant after they've already searched or browsed, not instead of doing so (Do people use AI assistants before or after searching?). In practice, assistants often act as a second step rather than the only source.
It's also worth asking how you would rank assistants fairly at all. The EBU study used human journalists, which is costly but credible. Other work in the collection shows that automated judging can drift badly: LLM-as-a-judge setups shifted their verdicts 31% of the time on complex tasks, while agent-based evaluators that gather evidence cut that to a fraction of a percent (Can agents evaluate AI outputs more reliably than language models?). Even human reviewers disagree with each other more than you might expect (Does AI theme-mapping perform as well as human reviewers?). So any "winner" in a news-accuracy league table carries real uncertainty, and the lead can change with each model update.
The takeaway: picking the best assistant is less useful than knowing that all of them currently misattribute sources often enough that clicking through to the original article is the habit that protects you. About 42% of chatbot news users already do this.
Sources 7 notes
A coordinated study of 22 public broadcasters in 18 countries had journalists evaluate over 3,000 AI responses on news topics. Nearly half contained significant errors, a third had serious sourcing problems, and a fifth showed major accuracy issues like hallucinations.
Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.
A 48-market survey found weekly chatbot news use rose from 7% to 10%, concentrated among under-35s and certain regions. Yet only 20% trust chatbot answers versus 37% for news overall, and 42% of users click through to original sources.
A 45-market survey finds AI chatbot use for news rose to 10% globally, with trust in chatbots correlating more strongly with use than trust in social media does. Among chatbot users, 44% trust news from them versus 17% of non-users, suggesting trust gates deliberate adoption.
A cross-surface panel study found assistant sessions come after search and browsing 20.6 percentage points more often than before, reversing the "answer engine" narrative. Assistant-only sessions are also more common than search-only sessions within the same users.
Show all 7 sources
Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.
UK government's Consult tool achieved F1 0.76 against expert reviewers, compared to F1 0.81 between two human reviewers. Differences rarely affected which themes ranked top, suggesting AI performance was competitive despite inherent subjectivity in theme assignment.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Digital News Report 2026
- Emerging uses of AI chatbots for news and what it means for journalism (Digital News Report 2026)
- Artificial intelligence is ineffective and potentially harmful for fact checking
- News Integrity in AI Assistants
- The Decision to Verify: How Warmth and User Characteristics Shape Reliance on Conversational Agents for Information Search
- How AI Is Changing Search Behaviors
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
- When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs