Writing

How to measure an AI support assistant

A deflection rate is a definition before it is a number, and the flattering definition is always available. The four metrics worth tracking, and the containment trade-off a randomised trial actually measured.

·11 min read

Nearly every number in this category is a definition wearing a percentage sign. Before you compare your deflection rate with anyone else's — or with the one on your own dashboard last quarter — it is worth working out which of several defensible definitions you are using.

Three honest answers, one inbox

Take a single month and compute the deflection rate three ways. All three are defensible. All three are in use commercially. They differ by more than most people's year-on-year improvement.

1,000
82%
23%
9%
“AI handled”
82.0%
Every conversation it replied in. The number on the slide.
“Resolved” by silence
63.1%
Nobody followed up, so it is counted as solved. Includes everyone who gave up.
Plausibly resolved
55.7%
Answered, no follow-up, and no sign they abandoned it.
A 26.3% spread, from one inbox74 people left without a word and 189 asked again. The middle bar counts the first group as a success, which is why a vendor whose billing depends on silence will quote it.
Fig. 01The same month under three definitions of deflection— raise the share who give up without saying anything →↳ 18% escalating is held constant; only the definition changes between the bars

The middle bar is the one that gets quoted, and it is the one with the conflict of interest baked in. If a vendor bills per resolution and defines a resolution as the absence of a follow-up, then every customer who gives up in silence is both a success and an invoice. We go through whose pricing works that way in the Fin alternatives comparison.

Any metric that treats someone leaving without a word as a success will improve fastest when your assistant gets worse.

What a real number looks like

There is no independent benchmark for AI deflection, and it is worth being blunt about that rather than quoting one of the circulating figures. Gartner (2024) measures self-service resolution at 14% — all self-service, surveyed in December 2023, not LLM agents. Vendor claims run from 50% to 80% on definitions they chose.

The most usable figure we found is Gorgias's platform data (2026): across a thousand-odd ecommerce brands, the median self-serves 45% of AI-touched tickets end to end and the top quartile 65%. It is a vendor publishing its own telemetry, so discount accordingly — but a median and a quartile are far more honest than an "up to" number, because they admit a distribution exists.

The trade you are actually making

Pushing containment up is not free, and for once there is a proper experiment rather than a case study. Researchers ran a randomised trial across 680,676 chats and 647 agents, comparing AI-eligible conversations with human-handled controls.

45%
Projected rating, out of 53.18Down from 3.57
Shorter conversations29%Of the contained share
Repeat contactsunchangedStatistically indistinguishable in the trial
The trade is worth making hereContainment buys speed and costs satisfaction, roughly in proportion. Repeat contacts did not rise — the authors read the ratings drop as a reaction to how the AI talks, not to whether the problem got solved. Endpoints are measured; everything between them is a straight line we drew.
Fig. 02Containment against satisfaction, and why late handoff is worst— raise containment, then switch when the handoff happens →↳ endpoints are measured in the trial; everything between them is linear interpolation we drew

The headline result (2026): conversations the AI handled with no human involvement at all were 64.6% shorter and rated 0.858 points lower out of five, from a baseline of 3.57. Crucially, repeat contacts were statistically unchanged — so people were not coming back more often. The authors read the ratings drop as a reaction to how the AI communicates rather than to whether it solved anything.

The more actionable finding is the second toggle. When the handoff came only after the customer was already annoyed, conversations ran 40.8% longer, repeat contacts rose six points, and ratings fell furthest of all. Holding on too long is worse than either extreme — which is the strongest argument available for making the exit easy to find rather than a reward for persistence.

The four metrics worth a dashboard

  1. Unanswered questions, as a list. Not a count — the actual text. This is the only metric that tells you what to do next, and it is the input to deciding what to train on.
  2. Plausible resolution rate — the strict definition from Figure 01. Track its trend, not its level, and never compare it to a vendor's.
  3. Time to a human, at the 90th percentile. Not the median. Gorgias found a median wait of ten hours after handoff and a 90th percentile of seventy-one, with 33% of AI-touched tickets never getting a human reply (2026). The tail is the customer experience.
  4. Satisfaction on contained conversations only. Blended CSAT hides the trade-off the figure above measures. Split it or you are averaging away the thing you need to see.

One caveat about the whole exercise

A good share of your customers would rather you had not done any of this. Gartner (2024) found 64% would prefer companies did not use AI for service and 53% would consider switching to a competitor over it — with the top stated concern being that it will get harder to reach a person. That is stated preference rather than measured behaviour, and people routinely use things they say they dislike. But it sets the bar: the assistant has to be better than the queue it replaced, not merely cheaper.

The counter-evidence is worth holding at the same time. The best-identified study of AI assisting humans — Brynjolfsson, Li and Raymond in the QJE (2025), across 5,179 agents — found a 14% rise in issues resolved per hour, 34% for the least experienced, with customer sentiment improving rather than falling. The difference between that result and the containment trial is not the technology. It is whether a person was still in the conversation.

If you want the cost side of the same ledger, the cost-per-conversation calculator takes your own volumes, and Zinx Signal makes the same case about measurement for landing pages that this post makes for support.

Common questions

What is a good deflection rate for an AI chatbot?
There is no independent benchmark. Gartner measures self-service resolution at 14%, which is a different thing. Vendor claims cluster between 50% and 80% using definitions they chose themselves, and Gorgias platform data puts the median ecommerce brand at 45% of AI-touched tickets self-served end to end, with the top quartile at 65%.
How should I calculate deflection rate honestly?
Count only conversations where the assistant answered, nobody asked again, and there is no sign the person abandoned the chat. Then publish that number next to the looser ones rather than instead of them. The gap between definitions is usually 20 points or more from the same month of the same inbox.
Does automating more conversations hurt customer satisfaction?
Yes, measurably. A randomised trial over 680,676 chats found conversations the AI finished with no human at all were 64.6% shorter but rated 0.858 points lower out of five. Repeat contacts did not rise, so the authors read it as a reaction to how the AI communicates rather than to whether the problem got solved.
What is the most useful metric for a support assistant?
The list of questions it could not answer. Deflection tells you how the assistant performed against content you already had; the unanswered list tells you what to write next, which is the only metric that compounds.
metricssupportguide

Try it on your own content

25 free replies, no card. Point it at your site, ask it the three questions you answer most, and read what it says.

Read next

retrieval·10 min read·2 figures to try

Why an AI assistant makes things up, and what retrieval fixes

Retrieval cuts invented answers by more than half and does not eliminate them. What the measured numbers are, why a confident wrong answer is the dangerous kind, and what to do about the remainder.

shopify·10 min read·2 figures to try

An AI assistant on a Shopify store: what it can and cannot see

Order status is 18% of ecommerce questions and answering it is a lookup. Here is what an assistant can read from your catalogue, what it must never read, and what the dull questions are worth.

support·8 min read·2 figures to try

Your AI assistant needs a way to reach a person

87% of customers say a company using AI for support must still offer a human. What that means in practice: the three endings a stuck conversation can have, and what each one costs you.

Written by Shardul Gautam, who builds Zinx Chat.