How often is AI wrong?
The most rigorous study to date found a significant problem in 45% of AI answers about the news. Here is what that number covers, what it does not, and why a single accuracy figure is misleading.
There is no single number, and anyone quoting one without saying what was measured is selling something. But there is one study large enough and independent enough to anchor the question.
The best figure available: 45%
In October 2025 the European Broadcasting Union and the BBC published the largest study of its kind. Twenty-two public service media organisations across 18 countries and 14 languages assessed more than 3,000 responses from ChatGPT, Copilot, Gemini and Perplexity to questions about current news.
- 45% of all answers had at least one significant issue.
- 20% had major accuracy problems, hallucinated details, outdated information presented as current.
- 31% had a sourcing problem, including attributing claims to outlets that never made them.
- 81% had at least one issue of any severity.
The researchers noted the pattern held across languages and territories, which is the finding that matters most. This is not one model having a bad month in one market. It is a consistent property of the category.
An earlier BBC-only study in February 2025 found 51% of responses to BBC-related news queries had significant issues. Over roughly eight months, on a comparable measure, the number moved from about half to a bit under half.
Why the number is not “the accuracy of AI”
Four caveats, all of which cut in different directions:
News is a hard case. It rewards recency, precise attribution and knowing what has changed since training, the model’s weakest ground. On stable, well-documented topics, error rates are lower.
News is also an extremely common case. Hundreds of millions of people ask these systems what is happening. A task being hard is not a reason to discount the failure rate on it.
“Significant issue” is not the same as “wrong.” The category includes sourcing failures where the underlying fact was right. A brief attributed to the wrong outlet is a real problem, but a different kind from an invented fact.
Rates vary sharply by model and by task. Gemini’s 72% sourcing-issue rate against under 25% for the others is a four-fold spread inside a single study. Any average across models hides more than it shows.
Why the number does not simply fall over time
It is reasonable to assume this improves with each release. Partly it does. But the 2025 OpenAI paper Why Language Models Hallucinate identifies a structural reason it does not improve as fast as capability does.
Most benchmarks score models on the percentage of questions answered correctly. Under that scoring, an honest “I don’t know” scores identically to a wrong answer, zero, while a confident guess is sometimes right. The training signal therefore favours guessing over abstaining, and it does so no matter how capable the underlying model becomes.
In other words: a large share of the residual error rate is not a capability limit. It is what the industry currently rewards.
The practical version
For anything with a consequence (a citation, a dosage, a policy, a date, a quotation, a number you are about to publish) assume the answer is unverified and treat verification as part of the task rather than an optional extra.
That advice sounds tediously conservative until you look at what the 45% is made of. It is not exotic failures on obscure questions. It is ordinary answers to ordinary questions, delivered in exactly the same confident register as the correct ones.