Which model should an AI support assistant run on?
Which model to run a grounded support assistant on, with prices read off each vendor's own page: the cheapest one that calls a tool reliably, what forced reasoning costs, and why retrieval matters more.
Almost every comparison of models asks which one is cleverest. For a support assistant that is the wrong question, because the model is not the part doing the hard work. Here is the right question, and what the answer costs.
The model is not the expensive part
A grounded support answer is a quotation with manners. The passages have already been found; the model's job is to read four of them and write two sentences that do not contradict any of them. That is not a task that rewards brilliance, and the price range on offer for doing it is extraordinary.
- GPT-6 Luna
- $1.40
- GPT-5.6 Luna
- $3.05
- Gemini 3.1 Flash-Lite
- $3.81
- Gemini 3.8 Flash
- $10
- Claude Haiku 4.5
- $14
- GPT-6 Sol
- $28
- GPT-6.1 Sol
- $28
- Claude Sonnet 5.5
- $28
- Claude Opus 5.5
- $56
- GPT-6 Astra
- $140
A hundredfold spread, for identical input and identical output length. Nothing further up that list improves your sources — it only improves the prose that quotes them. If your assistant is getting answers wrong, moving up this ladder is the most expensive way to not fix it.
You are not buying intelligence. You are buying a competent writer for a paragraph somebody else already researched.
What actually disqualifies a model
Two capabilities are non-negotiable, and benchmark tables rarely mention either.
- It has to call a function, reliably. The assistant decides what to search for, searches, and answers from what came back. A model that cannot be trusted to make that call is not a cheaper option, it is a broken one. Watch the endpoint as well as the model — on OpenAI's newest family, function calling is restricted or absent on the older of its two APIs, so the same model id behaves differently depending on how you reach it.
- It has to stream. A support answer that arrives all at once after four seconds feels broken; the same answer streaming from four hundred milliseconds feels immediate. This is the single largest perceived-quality lever, and it is free.
- It has to still exist next quarter. Read the retirement date before the benchmark. Models are switched off on published dates, and a retired model does not degrade gracefully — it returns errors, and every error is a customer told the assistant is unavailable.
The reasoning tax
Newer models think before answering, and most of them now do it by default. For a support assistant that is usually money spent on deliberation nobody asked for: the customer waits longer, the thinking is discarded before it reaches them, and the tokens are billed as output either way.
The part worth knowing before you choose is that you cannot always turn it off. Each vendor spells the off switch differently, several of their newest models refuse it outright, and on one lineup the documented floor is explicitly not a guarantee of zero.
- Input
- $0.00026
- The answer
- $0.00021
- Thinking, discarded before sending
- $0.00015
Thinking is worth paying for on a question that genuinely needs working through — a policy with conditions, a comparison that spans several documents. It is worth refusing on "what are your opening hours?". Since most support questions are the second kind, reasoning belongs behind a setting somebody chose rather than switched on for everybody.
The shortlist, by what you care about
There is no best model, only a best trade. Four priorities, and the model each one points at.
- Model
- GPT-6 Luna
- Vendor
- OpenAI
- Per million tokens
- $0.1 in, $0.5 out
- One grounded answer
- $0.00047
- Reasoning
- Can be switched off
- Context window
- 1,050K tokens
Prices come from OpenAI's pricing page, Anthropic's and Google's, each read on 2026-10-07. Two of them carry published end dates, which is worth knowing before a margin gets built on one.
The choice that matters more than any of these
Everything above moves the bill. What moves the answers is whether the right paragraph was in front of the model at all.
A strong model handed the wrong passage writes a confident, fluent, wrong answer — which is worse than an honest refusal, because it reads as authoritative and nobody checks it. A cheap model handed the right passage quotes it correctly. That asymmetry is the whole reason retrieval rather than a bigger model is what fixes invented answers, and why what you train it on beats what you run it on.
There is a second model in a grounded assistant, and almost nobody chooses it deliberately: the embedding model, which turns your documents and the customer's question into the vectors that decide what gets retrieved. It is priced in pennies and it sets your ceiling.
- OpenAI 3-small
- #49
- OpenAI 3-large
- #35
- Gemini Embedding 2
- #24
- Qwen3 Embedding 8B
- #14
- Voyage 4 Large
- #1
- Per million tokens
- $0.02
- Vector width
- 1,536, truncatable
- Longest chunk it will read
- 8,192 tokens
- RTEB English
- Rank 49
- MTEB retrieval mean
- 53.48
Two things are worth taking from that. The cheap, obvious default is genuinely mid-table — there is real retrieval quality available above it. And the ranking does not follow the price: the best-scoring model in that list is also one of the cheapest, because it has open weights and no vendor's margin on top.
The reason this is not simply free to fix: a vector index is built for one model at one width, and changing either makes every passage you have already indexed unfindable rather than throwing an error. Switching embedding models means re-indexing everything, which is why it is worth choosing carefully once rather than casually twice.
So if you have a budget to spend on quality, spend it in this order: better source documents, then better retrieval, then a better model. Most people do it backwards, because the chat model is the only one of the three with a leaderboard anybody reads.
What to avoid
- Free tiers of hosted models. They rate-limit and return errors under any real traffic, and each error is a visitor told the assistant is unavailable — on a site whose owner never chose that.
- A model with no published price. If you cannot work out what an answer costs, you cannot work out what a plan costs, and somebody eventually eats the difference.
- The frontier model, as a default. Offer it as a setting for hard questions. Making it the default spends fifty times the money on lookups that a lookup model answers identically.
- Any model you have not read a retirement date for. This is the one that becomes an incident rather than an invoice.
For what all of this adds up to per conversation, the cost-per-conversation post puts model spend next to the platform fee and the vendors who charge per resolution. And once it is running, measuring it will tell you far more than any benchmark — including whether a cheaper model would have done.
Common questions
- What is the best AI model for customer support?
- The cheapest one that calls a search tool reliably. A support answer is quoted out of your own documents, so the model is writing a sentence from text it has been handed — not recalling anything. Paying a hundred times more buys a better writer of that sentence, which is the cheap half of the work.
- What should I look for in a model for a support chatbot?
- Two things, and neither is intelligence: it has to call a function reliably, so it can go and fetch the right passage, and it has to stream, so somebody watching a cursor sees words appear. A model that fails either is unusable regardless of how it scores on a benchmark.
- Do I need a reasoning model for customer support?
- Mostly no. Newer models reason by default, and on a support assistant that is a tax: the customer waits longer, the reasoning is thrown away before they see it, and the tokens bill as output. It earns its keep on a policy question with conditions, which is why it belongs behind a setting rather than on by default.
- Does the embedding model matter more than the chat model?
- Far more. The model writes the sentence; retrieval decides whether the right paragraph was in front of it. A strong model given the wrong passage produces a confident, well-written wrong answer — which is worse than an honest refusal, because it reads as authoritative.
- How often do AI model prices and models change?
- Check the vendor's deprecation page before you build on anything. Models are switched off on published dates, and a retired model does not get worse — it starts returning errors, and every one is a customer told the assistant is unavailable. Pin a model you have read a retirement date for.