Writing

What to train an AI assistant on, and what it does with it

An assistant is not trained on your site in the way people assume. Here is what retrieval actually hands the model, which pages earn their place, and why the dull documents beat the marketing ones.

·9 min read

“Train it on your website” is the phrase everybody uses, and it gives almost everybody the wrong mental model. Nothing is memorised. Each question fetches a handful of passages, and those passages are the entire world the answer gets written from — which changes what is worth feeding it.

What actually happens to a page

A page you add is fetched, stripped to its text, cut into passages of a few hundred words, and turned into vectors — coordinates that put passages about delivery near other passages about delivery. Nothing is sent to a model at this stage, and no model learns anything.

When a visitor asks something, the question goes through the same process, and the passages nearest it come back. Those few passages are put in front of the model along with the question. That is the whole trick, and three useful consequences fall out of it:

  • Deleting a source works immediately. Its passages go with it; there is no retraining and no residue.
  • A wrong price on a page is a wrong price in the chat. Retrieval is loyal, not clever.
  • Wording matters more than volume. A passage written in your words will not match a question asked in the customer's.
Passages retrieved
Delivery & returns0.89Standard UK delivery takes 2–4 working days…
Delivery & returns0.71Orders placed before 2pm are dispatched the same day…
FAQ: shipping0.64We ship Monday to Friday, excluding bank holidays…
Passages above the threshold
3
What the model is given
the question and these passages
It answers, and citesStandard delivery is 2–4 working days in the UK. Orders placed before 2pm go out the same day.
Fig. 01What retrieval hands the model— ask one of the two questions →↳ the second one fails on purpose — nothing is close enough
The model is not the part that knows things. The passages are. A better model writes a better sentence from the same facts — and invents the same nothing when there are none.

Which pages earn their place

The instinct is to start with the home page and the about page. They are the two pages that answer the least. Gartner's finding (2024) that 43% of self-service failures come down to the customer not finding relevant content is a content problem before it is a software one — and the content that is missing is almost always operational rather than promotional.

  • How long does delivery take? — nothing to answer from
  • Can I return it if it doesn’t fit? — nothing to answer from
  • Is it machine washable? — nothing to answer from
  • Which size am I if I’m 5'9"? — nothing to answer from
  • What do you actually sell?
  • Do you ship to Spain, and what’s the duty? — nothing to answer from
  • My order says delivered but isn’t. — nothing to answer from
1 of 7 answerable — 14%Marketing pages are written to persuade someone who has not decided. Almost nothing a buyer asks after deciding is on them.
Fig. 02Which questions each kind of source can answer— switch sources on and off →↳ turn on marketing pages alone first, then add the dull ones

In rough order of value per page:

  1. Delivery, returns and warranty. Three pages that answer a third of everything, and the three most people have not written properly.
  2. Product detail. Materials, sizing, compatibility, what is in the box. The specifics, not the adjectives.
  3. FAQ pairs you write yourself. For the questions no page answers because nobody thought to write them down — and for the ones your pages answer in the wrong words.
  4. Marketing pages. Last, and mostly so the assistant can describe what you sell.

What breaks, and how you find out

Three failures account for nearly all of it, and the dashboard is where each one shows up:

  • The page needs JavaScript to show its text. A crawler gets an empty shell. You will see a source that trained successfully with almost no words in it — which is why ours warns when a page comes back under eighty words, and why the crawler page explains what it can and cannot read.
  • The content exists in the wrong words. “Dispatch” on the page, “shipping” in the question. An FAQ pair fixes it in a line.
  • The answer is in a system, not a document. Order status, account details, anything personal. No amount of training solves this one; it is what handoff is for.

The list of questions that found nothing is the thing to read weekly. It is a content brief written by your customers, in their own words, sorted by how many of them asked — and it is the only reliable way to make the assistant better. Everything else is guessing at a model.

How much is enough

Twenty good pages beat two hundred indifferent ones. Every extra page is another candidate competing for the handful of slots a question gets, so a site padded with near-duplicate content retrieves worse, not better. Add the pages that answer things, read the unanswered list, write the gaps, and stop.

If you want to know whether a change actually helped rather than feeling like it did, that is a measurement problem, and it is the same one as testing anything else on a site. Our sister product Zinx Signal is built for exactly that argument, and its piece on how many visitors a test actually needs is worth reading before you draw a conclusion from a quiet week.

In Zinx Chat, sources live under Training: add a page, upload a file, paste text, or point the crawler at a site and let it read what robots.txt allows. Every source can be opened to see exactly which passages it produced, which is usually how you discover a page you trusted contains four sentences and a cookie banner. The free 25 replies are enough to find that out on your own site this afternoon.

Common questions

What content should I train my AI chatbot on?
Policy pages first — delivery, returns, warranty — then product pages with specs and sizing, then FAQ pairs written in your own words for the questions no page answers. Marketing pages answer very little of what a buyer asks after they have decided.
Does an AI chatbot memorise my website?
No. Retrieval splits your content into passages and finds the handful closest to each question, then gives only those to the model. The model has not memorised your site and holds nothing between conversations, which is why a page you delete stops being answerable immediately.
Why does my AI chatbot say it does not know?
Usually it means retrieval found nothing close enough to the question. Either the content is not there at all, or it is there in words no customer would use. The fix is a short FAQ pair in the customer's wording, not a bigger model.
trainingretrievalsupport

Try it on your own content

25 free replies, no card. Point it at your site, ask it the three questions you answer most, and read what it says.

Read next

metrics·11 min read·2 figures to try

How to measure an AI support assistant

A deflection rate is a definition before it is a number, and the flattering definition is always available. The four metrics worth tracking, and the containment trade-off a randomised trial actually measured.

retrieval·10 min read·2 figures to try

Why an AI assistant makes things up, and what retrieval fixes

Retrieval cuts invented answers by more than half and does not eliminate them. What the measured numbers are, why a confident wrong answer is the dangerous kind, and what to do about the remainder.

support·8 min read·2 figures to try

Your AI assistant needs a way to reach a person

87% of customers say a company using AI for support must still offer a human. What that means in practice: the three endings a stuck conversation can have, and what each one costs you.

Written by Shardul Gautam, who builds Zinx Chat.