Bring your own key: when BYOK is actually cheaper
Past a few thousand answers a month the model is most of the bill, and buying it directly beats buying it bundled. Work out your own break-even, and see where the tokens in one reply actually go.
Bundled model pricing is convenient and it has a crossover point. Below it you are paying a little extra for someone else to hold the API key; above it you are paying a lot extra. The only question is where your volume sits.
Where the money in one reply goes
Start with the anatomy, because it explains the whole economic shape. A grounded support answer is not an expensive piece of generation — it is a cheap piece of generation attached to an expensive piece of context.
- System prompt and tools
- 220
- Retrieved passages (4)
- 2,000
- Conversation so far
- 300
- The question itself
- 40
- The answer it writes
- 420
The question is forty tokens. The answer is four hundred. Everything else — usually 80% or more of what you pay — is the retrieved text that makes the answer checkable rather than invented. That is a cost worth paying; the alternative is a confident guess. But it means your bill is governed by retrieval settings, not by how chatty your customers are.
You are not buying answers. You are buying the context that makes an answer safe to send, and the answer comes almost free with it.
Your break-even
Here is the comparison with both sides costed. On our key, a reply spends the platform charge plus the model's tokens converted into replies. On yours, it spends the platform charge only — and the provider bills you directly for the tokens.
- On our key
- $199
- On your own key
- $80
Two things fall out of dragging that slider. The crossover moves a long way depending on the model — on a cheap one the platform fee is most of the cost and bundling stays competitive much longer. And the saving at realistic volumes is real but not dramatic, which is the honest answer most comparison pages avoid giving.
What you take on with the key
- A second invoice. Someone has to reconcile two bills instead of one, and the provider's is itemised by token rather than by anything a finance team recognises.
- Your own rate limits. If your key hits a provider limit mid-conversation, the assistant degrades and it is your account's problem to raise. Bundled capacity is somebody else's pager.
- Model choice, which is a real benefit. Your key, your model — including one we have not added, through a gateway. If you have negotiated rates or committed spend, BYOK is how you use them.
- Your data boundary moves. Prompts go to your account under your agreement with the provider, not ours. For some buyers that is the entire reason to do it, independent of price.
Which to pick
- Under a thousand answers a month: bundled, without thinking about it. The saving is a rounding error against an hour of your time.
- A few thousand, on a cheap model: still bundled, probably. Check the figure, but the platform fee dominates.
- A few thousand, on a strong model: your own key. The tokens are the bill and you can buy them cheaper than we can resell them.
- Any volume, with a data-residency or vendor requirement: your own key, regardless of price. This is not a cost decision.
And whichever you pick, the lever with the largest effect on the bill is not the key — it is how much context each answer drags in. Training it on fewer, better documents reduces cost and improves accuracy at the same time, which is a rare combination. For the wider comparison against per-resolution and per-message vendors, the cost-per-conversation post puts all of them on one axis.
Common questions
- What does BYOK mean for an AI chat assistant?
- Bring your own key: you supply an API key from OpenAI, OpenRouter or a gateway, and the assistant calls the model on your account instead of ours. You get the provider's invoice directly and we charge only for the platform — retrieval, storage, streaming and the dashboard.
- At what volume does BYOK become cheaper?
- Usually past a few thousand answers a month, and it depends more on the model than the volume. On a cheap model the platform fee dominates and bundling is simpler; on a strong reasoning model the tokens are most of the cost and buying them directly wins. The calculator in this post runs your numbers through the same functions our server bills with.
- What are the downsides of bringing your own key?
- Two things worth pricing: a second invoice to reconcile, and your own rate limits to manage. If your key hits a provider limit mid-conversation, that is now your outage. Below a few thousand replies a month the saving rarely covers the administration.
- What makes up the token cost of one AI support reply?
- Nearly all of it is retrieved context. In a typical grounded reply the question is around 40 tokens and the answer around 420, while the passages fetched to make the answer checkable run into thousands. That is why retrieving four good passages beats eight mediocre ones on both cost and accuracy.