Can you trust ChatGPT's answers?
A practical answer: trust it by category, not in general. Here is where the failure rate is measurably high, where it is low, and the checks worth doing.
Trust it the way you would trust a fast, articulate, widely-read colleague who has never once said “I’m not sure.” The output is often excellent. The absence of a reliability signal is the problem.
That is not a rhetorical framing. It is what the measurements show.
Where the failure rate is measurably high
Current events and news. The 2025 European Broadcasting Union / BBC study (3,000+ responses, 22 media organisations, 18 countries) found 45% of AI answers about the news had at least one significant issue, and 20% had major accuracy problems. Sourcing was worse than facts: 31% misattributed material, sometimes to outlets that never published it.
Citations and references. This is the single worst category, because the format is learnable and the content is not. Legal citations have produced sanctions in over a thousand documented court proceedings, tracked in a public database maintained by researcher Damien Charlotin. The same failure mode applies to academic references, statute numbers and page citations.
Anything about a specific named entity that isn’t famous. Small companies, local officials, niche products. There is enough training data for the model to produce a confident-sounding profile, and not enough for it to be right.
Numbers you intend to publish. Statistics are the easiest thing in the world for a model to generate plausibly and the hardest for a reader to challenge.
Where it is genuinely reliable
Transforming text you supplied. Summarising, rewriting, changing tone, extracting structure from a document you pasted. The source of truth is in the prompt, so there is nothing to fabricate.
Well-documented, stable technical material. Standard library syntax, common algorithms, established procedures. Massively represented in training data and unchanged for years.
Explaining a concept you will then verify. As a way in to an unfamiliar subject, it is excellent, provided the exit is a real source.
Drafting and structuring. Where you are the domain expert and the model is the typist, the failure mode largely disappears.
The checks that actually work
Ask for sources, then open them. Asking alone does nothing, fabricated citations come with fabricated sources. The check is the click, not the request.
Ask the same question in a fresh conversation. Hallucinations are often unstable in a way that facts are not. Two independent runs diverging on a detail is a strong signal.
Watch for suspicious specificity. A model that produces an exact figure, a precise date and a named study for a question you would expect to be contested is doing pattern-completion, not recall.
Never accept a quotation. Attributed quotes are a high-risk category and always checkable.
The thing that will not save you
Confidence, fluency and formatting carry no information about accuracy. A fabricated case citation arrives with the same conviction as a real one, correct reporter format, plausible docket number, real judge’s name attached. There is no tell in the prose, and there will not be one, because well-formed output is what the system optimises for.
Human liars leak signals. This does not.
A one-line policy
If being wrong would cost you money, credibility, a client or a court appearance, the answer is a starting point and not a source. If being wrong costs you nothing, use it freely.
Most people already apply that rule to strangers on the internet. It transfers cleanly.