Writing

Which model should an AI support assistant run on?

Which model to run a grounded support assistant on, with prices read off each vendor's own page: the cheapest one that calls a tool reliably, what forced reasoning costs, and why retrieval matters more.

·11 min read

Almost every comparison of models asks which one is cleverest. For a support assistant that is the wrong question, because the model is not the part doing the hard work. Here is the right question, and what the answer costs.

The model is not the expensive part

A grounded support answer is a quotation with manners. The passages have already been found; the model's job is to read four of them and write two sentences that do not contradict any of them. That is not a task that rewards brilliance, and the price range on offer for doing it is extraordinary.

3,000
4 passages
GPT-6 Luna
$1.40
OpenAI — $0.00047 an answer
GPT-5.6 Luna
$3.05
OpenAI — $0.00102 an answer
Gemini 3.1 Flash-Lite
$3.81
Google — $0.00127 an answer
Gemini 3.8 Flash
$10
Google — $0.00349 an answer
Claude Haiku 4.5
$14
Anthropic — $0.00466 an answer
GPT-6 Sol
$28
OpenAI — $0.00932 an answer
GPT-6.1 Sol
$28
OpenAI — $0.00932 an answer
Claude Sonnet 5.5
$28
Anthropic — $0.00932 an answer
Claude Opus 5.5
$56
Anthropic — $0.019 an answer
GPT-6 Astra
$140
OpenAI — $0.047 an answer
Tokens per answer2,9802,560 in, 420 out
Cheapest, a month$1.40GPT-6 Luna
Spread, cheapest to dearest100×The same questions, the same passages
The model choice moves this bill 100-foldEvery row answers the same question from the same retrieved passages. Nothing above buys you better sources — only a better writer of the sentence that quotes them, which is the cheap half of the job.
Fig. 01The same answer, priced across ten models— drag your volume, then your retrieval depth →↳ every price read off the vendor's own pricing page on 2026-10-07; the token shape is measured on real replies

A hundredfold spread, for identical input and identical output length. Nothing further up that list improves your sources — it only improves the prose that quotes them. If your assistant is getting answers wrong, moving up this ladder is the most expensive way to not fix it.

You are not buying intelligence. You are buying a competent writer for a paragraph somebody else already researched.

What actually disqualifies a model

Two capabilities are non-negotiable, and benchmark tables rarely mention either.

  • It has to call a function, reliably. The assistant decides what to search for, searches, and answers from what came back. A model that cannot be trusted to make that call is not a cheaper option, it is a broken one. Watch the endpoint as well as the model — on OpenAI's newest family, function calling is restricted or absent on the older of its two APIs, so the same model id behaves differently depending on how you reach it.
  • It has to stream. A support answer that arrives all at once after four seconds feels broken; the same answer streaming from four hundred milliseconds feels immediate. This is the single largest perceived-quality lever, and it is free.
  • It has to still exist next quarter. Read the retirement date before the benchmark. Models are switched off on published dates, and a retired model does not degrade gracefully — it returns errors, and every error is a customer told the assistant is unavailable.

The reasoning tax

Newer models think before answering, and most of them now do it by default. For a support assistant that is usually money spent on deliberation nobody asked for: the customer waits longer, the thinking is discarded before it reaches them, and the tokens are billed as output either way.

The part worth knowing before you choose is that you cannot always turn it off. Each vendor spells the off switch differently, several of their newest models refuse it outright, and on one lineup the documented floor is explicitly not a guarantee of zero.

Model
300 tokens
Input
$0.00026
The retrieved passages and the question
The answer
$0.00021
420 tokens the customer actually reads
Thinking, discarded before sending
$0.00015
Avoidable — this model can be told not to
Cheapest reasoning settingNoneCan be switched off
Share of the bill that is thinking24%Tokens the customer never sees, billed as output
Here, thinking is a choice you can declineReasoning can be switched off outright, so a question that only needs a passage quoted back pays nothing for deliberation it did not need.
Fig. 02What thinking costs, and whether you can decline it— pick a model, then set how much it thinks →↳ the thinking-token count is yours to set — it varies by question, and we have no honest single figure for it

Thinking is worth paying for on a question that genuinely needs working through — a policy with conditions, a comparison that spans several documents. It is worth refusing on "what are your opening hours?". Since most support questions are the second kind, reasoning belongs behind a setting somebody chose rather than switched on for everybody.

The shortlist, by what you care about

There is no best model, only a best trade. Four priorities, and the model each one points at.

What matters most to you
Model
GPT-6 Luna
Vendor
OpenAI
Per million tokens
$0.1 in, $0.5 out
One grounded answer
$0.00047
Reasoning
Can be switched off
Context window
1,050K tokens
GPT-6 Luna, and whyA hundredth of the dearest model on this list, with the same retrieved passages going in. For a lookup answered from your own documents, the sentence costs almost nothing to write.
Fig. 03The pick, by priority— pick what matters most →↳ all four support function calling and streaming; the figure above prices them against each other

Prices come from OpenAI's pricing page, Anthropic's and Google's, each read on 2026-10-07. Two of them carry published end dates, which is worth knowing before a margin gets built on one.

The choice that matters more than any of these

Everything above moves the bill. What moves the answers is whether the right paragraph was in front of the model at all.

A strong model handed the wrong passage writes a confident, fluent, wrong answer — which is worse than an honest refusal, because it reads as authoritative and nobody checks it. A cheap model handed the right passage quotes it correctly. That asymmetry is the whole reason retrieval rather than a bigger model is what fixes invented answers, and why what you train it on beats what you run it on.

There is a second model in a grounded assistant, and almost nobody chooses it deliberately: the embedding model, which turns your documents and the customer's question into the vectors that decide what gets retrieved. It is priced in pennies and it sets your ceiling.

Embedding model
OpenAI 3-small
#49
Rank 49 on RTEB English — $0.02 per million tokens
OpenAI 3-large
#35
Rank 35 on RTEB English — $0.13 per million tokens
Gemini Embedding 2
#24
Rank 24 on RTEB English — $0.2 per million tokens
Qwen3 Embedding 8B
#14
Rank 14 on RTEB English — $0.01 per million tokens
Voyage 4 Large
#1
Rank 1 on RTEB English — $0.12 per million tokens
Per million tokens
$0.02
Vector width
1,536, truncatable
Longest chunk it will read
8,192 tokens
RTEB English
Rank 49
MTEB retrieval mean
53.48
Cheap, ubiquitous and distinctly mid-tableCheap, ubiquitous and distinctly mid-table. The sensible default, and the first thing to improve once retrieval is the bottleneck.
Fig. 04The model that decides which paragraph gets quoted— switch embedding models →↳ ranks are positions on MTEB's current English retrieval board, which publishes ranks but withholds aggregate scores because half its tasks are private

Two things are worth taking from that. The cheap, obvious default is genuinely mid-table — there is real retrieval quality available above it. And the ranking does not follow the price: the best-scoring model in that list is also one of the cheapest, because it has open weights and no vendor's margin on top.

The reason this is not simply free to fix: a vector index is built for one model at one width, and changing either makes every passage you have already indexed unfindable rather than throwing an error. Switching embedding models means re-indexing everything, which is why it is worth choosing carefully once rather than casually twice.

So if you have a budget to spend on quality, spend it in this order: better source documents, then better retrieval, then a better model. Most people do it backwards, because the chat model is the only one of the three with a leaderboard anybody reads.

What to avoid

  1. Free tiers of hosted models. They rate-limit and return errors under any real traffic, and each error is a visitor told the assistant is unavailable — on a site whose owner never chose that.
  2. A model with no published price. If you cannot work out what an answer costs, you cannot work out what a plan costs, and somebody eventually eats the difference.
  3. The frontier model, as a default. Offer it as a setting for hard questions. Making it the default spends fifty times the money on lookups that a lookup model answers identically.
  4. Any model you have not read a retirement date for. This is the one that becomes an incident rather than an invoice.

For what all of this adds up to per conversation, the cost-per-conversation post puts model spend next to the platform fee and the vendors who charge per resolution. And once it is running, measuring it will tell you far more than any benchmark — including whether a cheaper model would have done.

Common questions

What is the best AI model for customer support?
The cheapest one that calls a search tool reliably. A support answer is quoted out of your own documents, so the model is writing a sentence from text it has been handed — not recalling anything. Paying a hundred times more buys a better writer of that sentence, which is the cheap half of the work.
What should I look for in a model for a support chatbot?
Two things, and neither is intelligence: it has to call a function reliably, so it can go and fetch the right passage, and it has to stream, so somebody watching a cursor sees words appear. A model that fails either is unusable regardless of how it scores on a benchmark.
Do I need a reasoning model for customer support?
Mostly no. Newer models reason by default, and on a support assistant that is a tax: the customer waits longer, the reasoning is thrown away before they see it, and the tokens bill as output. It earns its keep on a policy question with conditions, which is why it belongs behind a setting rather than on by default.
Does the embedding model matter more than the chat model?
Far more. The model writes the sentence; retrieval decides whether the right paragraph was in front of it. A strong model given the wrong passage produces a confident, well-written wrong answer — which is worse than an honest refusal, because it reads as authoritative.
How often do AI model prices and models change?
Check the vendor's deprecation page before you build on anything. Models are switched off on published dates, and a retired model does not get worse — it starts returning errors, and every one is a customer told the assistant is unavailable. Pin a model you have read a retirement date for.
modelspricingretrieval

Try it on your own content

25 free replies, no card. Point it at your site, ask it the three questions you answer most, and read what it says.

Read next

retrieval·10 min read·2 figures to try

Why an AI assistant makes things up, and what retrieval fixes

Retrieval cuts invented answers by more than half and does not eliminate them. What the measured numbers are, why a confident wrong answer is the dangerous kind, and what to do about the remainder.

pricing·8 min read·3 figures to try

Bring your own key: which provider, and when it pays

Bring a key from OpenAI, OpenRouter or Vercel AI Gateway. What each one gets you, where the tokens in a reply actually go, and the volume at which buying the model directly beats buying it bundled.

training·9 min read·2 figures to try

What to train an AI assistant on, and what it does with it

An assistant is not trained on your site in the way people assume. Here is what retrieval actually hands the model, which pages earn their place, and why the dull documents beat the marketing ones.

Written by Shardul Gautam, who builds Zinx Chat.