How to measure an AI support assistant
A deflection rate is a definition before it is a number, and the flattering definition is always available. The four metrics worth tracking, and the containment trade-off a randomised trial actually measured.
Nearly every number in this category is a definition wearing a percentage sign. Before you compare your deflection rate with anyone else's — or with the one on your own dashboard last quarter — it is worth working out which of several defensible definitions you are using.
Three honest answers, one inbox
Take a single month and compute the deflection rate three ways. All three are defensible. All three are in use commercially. They differ by more than most people's year-on-year improvement.
- “AI handled”
- 82.0%
- “Resolved” by silence
- 63.1%
- Plausibly resolved
- 55.7%
The middle bar is the one that gets quoted, and it is the one with the conflict of interest baked in. If a vendor bills per resolution and defines a resolution as the absence of a follow-up, then every customer who gives up in silence is both a success and an invoice. We go through whose pricing works that way in the Fin alternatives comparison.
Any metric that treats someone leaving without a word as a success will improve fastest when your assistant gets worse.
What a real number looks like
There is no independent benchmark for AI deflection, and it is worth being blunt about that rather than quoting one of the circulating figures. Gartner (2024) measures self-service resolution at 14% — all self-service, surveyed in December 2023, not LLM agents. Vendor claims run from 50% to 80% on definitions they chose.
The most usable figure we found is Gorgias's platform data (2026): across a thousand-odd ecommerce brands, the median self-serves 45% of AI-touched tickets end to end and the top quartile 65%. It is a vendor publishing its own telemetry, so discount accordingly — but a median and a quartile are far more honest than an "up to" number, because they admit a distribution exists.
The trade you are actually making
Pushing containment up is not free, and for once there is a proper experiment rather than a case study. Researchers ran a randomised trial across 680,676 chats and 647 agents, comparing AI-eligible conversations with human-handled controls.
The headline result (2026): conversations the AI handled with no human involvement at all were 64.6% shorter and rated 0.858 points lower out of five, from a baseline of 3.57. Crucially, repeat contacts were statistically unchanged — so people were not coming back more often. The authors read the ratings drop as a reaction to how the AI communicates rather than to whether it solved anything.
The more actionable finding is the second toggle. When the handoff came only after the customer was already annoyed, conversations ran 40.8% longer, repeat contacts rose six points, and ratings fell furthest of all. Holding on too long is worse than either extreme — which is the strongest argument available for making the exit easy to find rather than a reward for persistence.
The four metrics worth a dashboard
- Unanswered questions, as a list. Not a count — the actual text. This is the only metric that tells you what to do next, and it is the input to deciding what to train on.
- Plausible resolution rate — the strict definition from Figure 01. Track its trend, not its level, and never compare it to a vendor's.
- Time to a human, at the 90th percentile. Not the median. Gorgias found a median wait of ten hours after handoff and a 90th percentile of seventy-one, with 33% of AI-touched tickets never getting a human reply (2026). The tail is the customer experience.
- Satisfaction on contained conversations only. Blended CSAT hides the trade-off the figure above measures. Split it or you are averaging away the thing you need to see.
One caveat about the whole exercise
A good share of your customers would rather you had not done any of this. Gartner (2024) found 64% would prefer companies did not use AI for service and 53% would consider switching to a competitor over it — with the top stated concern being that it will get harder to reach a person. That is stated preference rather than measured behaviour, and people routinely use things they say they dislike. But it sets the bar: the assistant has to be better than the queue it replaced, not merely cheaper.
The counter-evidence is worth holding at the same time. The best-identified study of AI assisting humans — Brynjolfsson, Li and Raymond in the QJE (2025), across 5,179 agents — found a 14% rise in issues resolved per hour, 34% for the least experienced, with customer sentiment improving rather than falling. The difference between that result and the containment trial is not the technology. It is whether a person was still in the conversation.
If you want the cost side of the same ledger, the cost-per-conversation calculator takes your own volumes, and Zinx Signal makes the same case about measurement for landing pages that this post makes for support.
Common questions
- What is a good deflection rate for an AI chatbot?
- There is no independent benchmark. Gartner measures self-service resolution at 14%, which is a different thing. Vendor claims cluster between 50% and 80% using definitions they chose themselves, and Gorgias platform data puts the median ecommerce brand at 45% of AI-touched tickets self-served end to end, with the top quartile at 65%.
- How should I calculate deflection rate honestly?
- Count only conversations where the assistant answered, nobody asked again, and there is no sign the person abandoned the chat. Then publish that number next to the looser ones rather than instead of them. The gap between definitions is usually 20 points or more from the same month of the same inbox.
- Does automating more conversations hurt customer satisfaction?
- Yes, measurably. A randomised trial over 680,676 chats found conversations the AI finished with no human at all were 64.6% shorter but rated 0.858 points lower out of five. Repeat contacts did not rise, so the authors read it as a reaction to how the AI communicates rather than to whether the problem got solved.
- What is the most useful metric for a support assistant?
- The list of questions it could not answer. Deflection tells you how the assistant performed against content you already had; the unanswered list tells you what to write next, which is the only metric that compounds.