Writing

Bring your own key: when BYOK is actually cheaper

Past a few thousand answers a month the model is most of the bill, and buying it directly beats buying it bundled. Work out your own break-even, and see where the tokens in one reply actually go.

·8 min read

Bundled model pricing is convenient and it has a crossover point. Below it you are paying a little extra for someone else to hold the API key; above it you are paying a lot extra. The only question is where your volume sits.

Where the money in one reply goes

Start with the anatomy, because it explains the whole economic shape. A grounded support answer is not an expensive piece of generation — it is a cheap piece of generation attached to an expensive piece of context.

4 passages
Model
System prompt and tools
220
Retrieved passages (4)
2,000
Conversation so far
300
The question itself
40
The answer it writes
420
Charged at $0.50 per million, against $0.10 going in
Cost of one answer$0.000472,560 in, 420 out
Per 1,000 answers$0.47
Retrieved text78%Of the input, at this setting
Retrieval is the billThe question is forty tokens and the answer four hundred. Nearly everything you pay for is the context fetched to make the answer checkable — which is the argument for retrieving four good passages rather than eight mediocre ones.
Fig. 01The tokens in one grounded reply— change how many passages get retrieved →↳ token counts are ours, measured on real replies; prices are the provider's list prices

The question is forty tokens. The answer is four hundred. Everything else — usually 80% or more of what you pay — is the retrieved text that makes the answer checkable rather than invented. That is a cost worth paying; the alternative is a confident guess. But it means your bill is governed by retrieval settings, not by how chatty your customers are.

You are not buying answers. You are buying the context that makes an answer safe to send, and the answer comes almost free with it.

Your break-even

Here is the comparison with both sides costed. On our key, a reply spends the platform charge plus the model's tokens converted into replies. On yours, it spends the platform charge only — and the provider bills you directly for the tokens.

3,000
Model
On our key
$199
Business plan
On your own key
$80
Pro plan, plus $1.40 billed to you by the provider
Replies spent per answer0.40Platform only, on your key
On our key0.87Platform plus the model
Model cost per answer$0.00047At the provider's list price
Your own key saves about $119 a monthPast a few thousand answers the model is most of the bill, and buying it directly is cheaper than buying it bundled. Assumes 2,560 tokens in and 420 out per answer. A longer knowledge base does not change this; more retrieved passages do.
Fig. 02Bundled against bring-your-own, at your volume— drag your monthly volume, then switch models →↳ every figure runs through the same functions that bill a real workspace, so this cannot drift from an invoice

Two things fall out of dragging that slider. The crossover moves a long way depending on the model — on a cheap one the platform fee is most of the cost and bundling stays competitive much longer. And the saving at realistic volumes is real but not dramatic, which is the honest answer most comparison pages avoid giving.

What you take on with the key

  • A second invoice. Someone has to reconcile two bills instead of one, and the provider's is itemised by token rather than by anything a finance team recognises.
  • Your own rate limits. If your key hits a provider limit mid-conversation, the assistant degrades and it is your account's problem to raise. Bundled capacity is somebody else's pager.
  • Model choice, which is a real benefit. Your key, your model — including one we have not added, through a gateway. If you have negotiated rates or committed spend, BYOK is how you use them.
  • Your data boundary moves. Prompts go to your account under your agreement with the provider, not ours. For some buyers that is the entire reason to do it, independent of price.

Which to pick

  1. Under a thousand answers a month: bundled, without thinking about it. The saving is a rounding error against an hour of your time.
  2. A few thousand, on a cheap model: still bundled, probably. Check the figure, but the platform fee dominates.
  3. A few thousand, on a strong model: your own key. The tokens are the bill and you can buy them cheaper than we can resell them.
  4. Any volume, with a data-residency or vendor requirement: your own key, regardless of price. This is not a cost decision.

And whichever you pick, the lever with the largest effect on the bill is not the key — it is how much context each answer drags in. Training it on fewer, better documents reduces cost and improves accuracy at the same time, which is a rare combination. For the wider comparison against per-resolution and per-message vendors, the cost-per-conversation post puts all of them on one axis.

Common questions

What does BYOK mean for an AI chat assistant?
Bring your own key: you supply an API key from OpenAI, OpenRouter or a gateway, and the assistant calls the model on your account instead of ours. You get the provider's invoice directly and we charge only for the platform — retrieval, storage, streaming and the dashboard.
At what volume does BYOK become cheaper?
Usually past a few thousand answers a month, and it depends more on the model than the volume. On a cheap model the platform fee dominates and bundling is simpler; on a strong reasoning model the tokens are most of the cost and buying them directly wins. The calculator in this post runs your numbers through the same functions our server bills with.
What are the downsides of bringing your own key?
Two things worth pricing: a second invoice to reconcile, and your own rate limits to manage. If your key hits a provider limit mid-conversation, that is now your outage. Below a few thousand replies a month the saving rarely covers the administration.
What makes up the token cost of one AI support reply?
Nearly all of it is retrieved context. In a typical grounded reply the question is around 40 tokens and the answer around 420, while the passages fetched to make the answer checkable run into thousands. That is why retrieving four good passages beats eight mediocre ones on both cost and accuracy.
pricingbyokmodels

Try it on your own content

25 free replies, no card. Point it at your site, ask it the three questions you answer most, and read what it says.

Read next

pricing·9 min read·3 figures to try

What an AI support assistant actually costs per conversation

Worked out rather than asserted: what a thousand support conversations cost under per-resolution, per-message and per-reply pricing, with a calculator you can put your own numbers into.

pricing·11 min read·2 figures to try

Intercom Fin alternatives, and what each one bills you for

Fin charges $0.99 when nobody asks a follow-up question. Eight alternatives, compared on the event that triggers a charge rather than the headline price — because that is what decides the invoice.

pricing·10 min read·2 figures to try

Per resolution, per message, per reply: AI chat pricing compared

Intercom Fin, Zendesk, Gorgias, Tidio, Chatbase, SiteGPT, CustomGPT and Crisp bill for different events. Here is what each one charges for, what that does to your invoice, and when each model is the right one.

Written by Shardul Gautam, who builds Zinx Chat.