Writing

Why an AI assistant makes things up, and what retrieval fixes

Retrieval cuts invented answers by more than half and does not eliminate them. What the measured numbers are, why a confident wrong answer is the dangerous kind, and what to do about the remainder.

·10 min read

A language model does not look an answer up and then tell you. It predicts text that fits the question. Most of the time the most plausible text is also true, which is why this works at all — and it is also why the failures are fluent, specific and completely wrong.

What changes when you hand it the source

The cleanest way to understand grounding is to ask the same question three ways. A shipping question is ideal, because a wrong answer is both easy to generate and expensive to give.

What the model was given

A shopper asks: “Do you ship to Norway, and how long does it take?”

Yes, we ship to Norway. Delivery usually takes 3–5 business days and shipping is free over $50.

No source to show.

Not grounded in anythingFluent, specific, and invented. Nothing in it came from your store — the model is completing a sentence that sounds like a shipping policy.
Fig. 01The same question, with and without a source— switch what the model was given →↳ all three answers are plausible; only one of them is checkable

Look at the first answer carefully. "3–5 business days" and "free over $50" are not random — they are the most statistically ordinary shipping policy in the English language. That is what makes the failure dangerous: it reads exactly like a correct answer, because correct answers to this question usually look like that.

The middle option is the one most tools actually ship, and the subtlest failure. The source is right there, and the model still averaged a table it was shown whole. Retrieval is not just "give it the page" — which passage you hand it is most of the work.

Fluency and accuracy are unrelated properties. A model has no way of sounding less sure, so it is your job to give it something to be sure about.

How much retrieval actually buys

This is where the published research is more useful than any vendor claim, and less flattering.

The task
Best model measured
1.8%
A current OpenAI model
3.1%
A small open model
3.7%
A large open model
4.1%

The closest published proxy for retrieval: the document is in the prompt and the only job is to stay faithful to it. Even here, nothing reaches zero. — Vectara, 2026. Scale is 6% full width; the two tasks are not comparable to each other.

Fig. 02Measured error rates, with sources in the prompt— switch between an easy grounding task and a hard one →↳ the two panels are different studies measuring different tasks; the bars are not comparable across them

The strongest evidence for the direction comes from Stanford's RegLab and HAI (2024), who tested purpose-built legal research tools on more than 200 pre-registered queries. With retrieval, hallucination ran above 17% for two products and above 34% for a third. Without it, the same underlying model was wrong on 58–82% of the same questions. Retrieval cut the error rate by more than half — and left one answer in six wrong, in a domain where being wrong is professional negligence.

Even on the easy version of the task, nothing reaches zero. Google DeepMind's FACTS Grounding benchmark (2025) hands the model the document and asks only that it stay faithful to it; the best model was fully grounded in 83.6% of answers. Vectara's current leaderboard (2026) measures the narrowest version — summarise this document, do not contradict it — and the best models still land between 1.8% and 4%.

Citations help, but not the way you think

Showing sources is the obvious remedy, and the research on it has an uncomfortable finding. A study published at AAAI (2025) varied how many citations an answer carried and whether they were relevant or random. Trust rose significantly when citations were shown — even when the citations were random — and fell when participants actually clicked through and checked them.

So citations do two things at once. They let a careful reader verify, which is the point. And they act as a trust signal for everyone else, which means they can make a wrong answer more convincing. That is an argument for showing them and for never treating them as proof the answer is right.

What to actually do about the remainder

  • Retrieve less, better. Four precise passages beat eight vague ones. More context is not more grounding; it is more opportunity to average.
  • Make "I don't know" a first-class answer. Models are measurably bad at refusing unanswerable questions, so the refusal has to be designed rather than hoped for.
  • Give the dead end somewhere to go. An honest "I cannot answer that, here is who can" is a better outcome than a confident guess, and it is the difference between a lost customer and a conversation.
  • Treat your own content as untrusted. A product description is a field a supplier feed can write to, which makes it a prompt-injection vector. The system prompt is the second line of defence; the first is not giving the assistant data it must not reveal.
  • Read the unanswered questions. Most residual wrong answers are a missing page, not a broken model.

None of this is unique to support. The same discipline shows up in PixGuard's problem — deciding whether an image has been tampered with — where a confident answer with no evidence trail is worth less than an uncertain one that shows its working. If you want to see how we decide what goes into the context in the first place, that is the retrieval post, and the measurement guide covers how to tell whether any of it is working.

Common questions

Does retrieval-augmented generation stop hallucinations?
No. It reduces them substantially and leaves a real remainder. Stanford researchers found purpose-built legal tools with retrieval still hallucinated on more than one query in six, against 58-82% for the same model without retrieval. Anyone claiming retrieval eliminates hallucination is selling something.
Why does a chatbot answer confidently when it is wrong?
Because it is not retrieving a fact, it is predicting plausible text. A shipping policy it has never seen still has a predictable shape — a number of days, a threshold, a currency — so the model fills that shape in. Fluency and accuracy are separate properties, which is exactly why confident wrong answers are hard to spot.
Do citations make an AI answer more trustworthy?
Partly, and not in the way you would hope. A 2025 study found showing citations raised self-reported trust even when the citations were randomly chosen, and that trust fell when participants actually checked them. So citations help readers who verify and flatter the ones who do not — which makes them necessary, not sufficient.
How do I reduce wrong answers from my support assistant?
Narrow the retrieved context, prefer specific documents over marketing pages, and make the assistant say it does not know rather than guess. Then read the questions it could not answer and write the missing page — the remaining errors are mostly a documentation gap, not a model problem.
retrievalaccuracyguide

Try it on your own content

25 free replies, no card. Point it at your site, ask it the three questions you answer most, and read what it says.

Read next

metrics·11 min read·2 figures to try

How to measure an AI support assistant

A deflection rate is a definition before it is a number, and the flattering definition is always available. The four metrics worth tracking, and the containment trade-off a randomised trial actually measured.

shopify·10 min read·2 figures to try

An AI assistant on a Shopify store: what it can and cannot see

Order status is 18% of ecommerce questions and answering it is a lookup. Here is what an assistant can read from your catalogue, what it must never read, and what the dull questions are worth.

training·9 min read·2 figures to try

What to train an AI assistant on, and what it does with it

An assistant is not trained on your site in the way people assume. Here is what retrieval actually hands the model, which pages earn their place, and why the dull documents beat the marketing ones.

Written by Shardul Gautam, who builds Zinx Chat.