Why an AI assistant makes things up, and what retrieval fixes
Retrieval cuts invented answers by more than half and does not eliminate them. What the measured numbers are, why a confident wrong answer is the dangerous kind, and what to do about the remainder.
A language model does not look an answer up and then tell you. It predicts text that fits the question. Most of the time the most plausible text is also true, which is why this works at all — and it is also why the failures are fluent, specific and completely wrong.
What changes when you hand it the source
The cleanest way to understand grounding is to ask the same question three ways. A shipping question is ideal, because a wrong answer is both easy to generate and expensive to give.
A shopper asks: “Do you ship to Norway, and how long does it take?”
Yes, we ship to Norway. Delivery usually takes 3–5 business days and shipping is free over $50.
No source to show.
Look at the first answer carefully. "3–5 business days" and "free over $50" are not random — they are the most statistically ordinary shipping policy in the English language. That is what makes the failure dangerous: it reads exactly like a correct answer, because correct answers to this question usually look like that.
The middle option is the one most tools actually ship, and the subtlest failure. The source is right there, and the model still averaged a table it was shown whole. Retrieval is not just "give it the page" — which passage you hand it is most of the work.
Fluency and accuracy are unrelated properties. A model has no way of sounding less sure, so it is your job to give it something to be sure about.
How much retrieval actually buys
This is where the published research is more useful than any vendor claim, and less flattering.
- Best model measured
- 1.8%
- A current OpenAI model
- 3.1%
- A small open model
- 3.7%
- A large open model
- 4.1%
The closest published proxy for retrieval: the document is in the prompt and the only job is to stay faithful to it. Even here, nothing reaches zero. — Vectara, 2026. Scale is 6% full width; the two tasks are not comparable to each other.
The strongest evidence for the direction comes from Stanford's RegLab and HAI (2024), who tested purpose-built legal research tools on more than 200 pre-registered queries. With retrieval, hallucination ran above 17% for two products and above 34% for a third. Without it, the same underlying model was wrong on 58–82% of the same questions. Retrieval cut the error rate by more than half — and left one answer in six wrong, in a domain where being wrong is professional negligence.
Even on the easy version of the task, nothing reaches zero. Google DeepMind's FACTS Grounding benchmark (2025) hands the model the document and asks only that it stay faithful to it; the best model was fully grounded in 83.6% of answers. Vectara's current leaderboard (2026) measures the narrowest version — summarise this document, do not contradict it — and the best models still land between 1.8% and 4%.
Citations help, but not the way you think
Showing sources is the obvious remedy, and the research on it has an uncomfortable finding. A study published at AAAI (2025) varied how many citations an answer carried and whether they were relevant or random. Trust rose significantly when citations were shown — even when the citations were random — and fell when participants actually clicked through and checked them.
So citations do two things at once. They let a careful reader verify, which is the point. And they act as a trust signal for everyone else, which means they can make a wrong answer more convincing. That is an argument for showing them and for never treating them as proof the answer is right.
What to actually do about the remainder
- Retrieve less, better. Four precise passages beat eight vague ones. More context is not more grounding; it is more opportunity to average.
- Make "I don't know" a first-class answer. Models are measurably bad at refusing unanswerable questions, so the refusal has to be designed rather than hoped for.
- Give the dead end somewhere to go. An honest "I cannot answer that, here is who can" is a better outcome than a confident guess, and it is the difference between a lost customer and a conversation.
- Treat your own content as untrusted. A product description is a field a supplier feed can write to, which makes it a prompt-injection vector. The system prompt is the second line of defence; the first is not giving the assistant data it must not reveal.
- Read the unanswered questions. Most residual wrong answers are a missing page, not a broken model.
None of this is unique to support. The same discipline shows up in PixGuard's problem — deciding whether an image has been tampered with — where a confident answer with no evidence trail is worth less than an uncertain one that shows its working. If you want to see how we decide what goes into the context in the first place, that is the retrieval post, and the measurement guide covers how to tell whether any of it is working.
Common questions
- Does retrieval-augmented generation stop hallucinations?
- No. It reduces them substantially and leaves a real remainder. Stanford researchers found purpose-built legal tools with retrieval still hallucinated on more than one query in six, against 58-82% for the same model without retrieval. Anyone claiming retrieval eliminates hallucination is selling something.
- Why does a chatbot answer confidently when it is wrong?
- Because it is not retrieving a fact, it is predicting plausible text. A shipping policy it has never seen still has a predictable shape — a number of days, a threshold, a currency — so the model fills that shape in. Fluency and accuracy are separate properties, which is exactly why confident wrong answers are hard to spot.
- Do citations make an AI answer more trustworthy?
- Partly, and not in the way you would hope. A 2025 study found showing citations raised self-reported trust even when the citations were randomly chosen, and that trust fell when participants actually checked them. So citations help readers who verify and flatter the ones who do not — which makes them necessary, not sufficient.
- How do I reduce wrong answers from my support assistant?
- Narrow the retrieved context, prefer specific documents over marketing pages, and make the assistant say it does not know rather than guess. Then read the questions it could not answer and write the missing page — the remaining errors are mostly a documentation gap, not a model problem.