
Want to measure your chatbot’s ROI? Twelve metrics cover performance, efficiency and customer satisfaction. Here is the short version:
- Customer Satisfaction (CSAT): How happy users are with an interaction.
- Net Promoter Score (NPS): How likely users say they are to recommend you.
- Customer Effort Score (CES): How hard users had to work to get their issue resolved.
- Total Conversations: How often the chatbot gets used at all.
- Chat Duration: How long conversations run.
- Chat Exit Rate: How often users leave before resolution.
- Response Time: How fast the chatbot replies.
- Resolution Time: How long the whole issue takes to close.
- Self-Service Rate: Share of issues resolved without a human.
- Sales Success Rate: Share of sales conversations that convert.
- Sales Amount: Revenue attributable to chatbot interactions.
- Cost Per Chat: Total chatbot cost divided by conversations handled.
One thing to settle first. There are no industry-standard “good” values for most of these. An earlier version of this article published target ranges for chat duration, resolution time and CES as though they were established benchmarks. They were not sourced, and they have been removed. The reliable method is to measure your own baseline in the first month and judge every later month against it. A number pulled from an article — including this one — tells you nothing about your customers, your product or your query mix.
1. Customer Satisfaction (CSAT)
CSAT measures whether an interaction met the user’s expectation. It is the bluntest of the three satisfaction metrics and the easiest to collect.
How to Calculate CSAT:
- Run a post-chat survey asking users to rate the interaction, commonly on a 1-5 scale.
- Count the proportion rating 4 or 5 as satisfied.
- Formula: CSAT = (Number of satisfied customers ÷ Total number of respondents) × 100
Note the sampling problem before you trust the number. Only a fraction of users answer a post-chat survey, and people with strong feelings answer more often. A rising CSAT with a falling response rate may mean nothing at all, so track both.
What to Track:
- Overall satisfaction with the interaction.
- Accuracy of the chatbot’s responses.
- Ease of communication.
- Whether the issue was actually resolved.
Tips for Measuring CSAT Effectively:
- Keep surveys short.
- Ask immediately after the interaction ends.
- Include an optional comment box — the comments are usually more useful than the score.
- Track the trend, not the absolute value.
Ways to Improve CSAT:
- Keep the knowledge base current, so answers are right.
- Adjust conversation flows based on where users get stuck.
- Use sentiment analysis to catch dissatisfaction mid-conversation.
- Hand off to a human when the chatbot cannot resolve the issue — quickly, rather than after three failed attempts.
2. Net Promoter Score (NPS)
NPS asks how likely a user is to recommend you. It was introduced by Fred Reichheld in “The One Number You Need to Grow” (Harvard Business Review, December 2003), which is where the promoter, passive and detractor bands come from.
Treat it with some caution. The claim that NPS predicts growth better than other satisfaction measures has been contested since publication, and independent replications have not consistently reproduced it. NPS is a reasonable tracking metric. It is not a law of business.
How to Calculate NPS:
After an interaction, ask: “On a scale of 0-10, how likely are you to recommend our chatbot service to others?”
Group responses:
- Promoters (9-10)
- Passives (7-8)
- Detractors (0-6)
Formula: NPS = % of Promoters – % of Detractors
Because passives are discarded, two very different response distributions can produce the same score. Always look at the underlying breakdown, not just the headline number.
Key Metrics to Watch:
-
Response Breakdown
The percentages of promoters, passives and detractors, and how each moves. -
Follow-Up Feedback
- Ask promoters what specifically worked.
- Ask detractors what went wrong.
- Look for themes that repeat.
Tips for Using NPS Effectively:
- Ask at consistent points in the customer journey, or the trend line is meaningless.
- Follow up with detractors promptly, while they still remember the interaction.
- Measure NPS separately for sales, support and general enquiries — they behave differently.
How to Improve NPS:
- Improve accuracy by retraining on the queries that failed.
- Personalise based on user history where you have it.
- Make the handoff to a human fast and context-preserving.
- Fix the complaints that repeat, rather than the loudest one.
When to Measure NPS:
- Transactional: right after an interaction.
- Relationship-based: periodically, for overall brand perception.
Comparing the two shows whether a good individual experience is translating into anything durable. Often it does not.
Segment the data by user group, interaction type, time of day, chat duration and issue complexity before drawing conclusions from a single figure.
3. Customer Effort Score (CES)
CES measures how much work the customer had to do. It comes from Dixon, Freeman and Toman’s “Stop Trying to Delight Your Customers” (Harvard Business Review, July 2010), whose argument was that reducing effort does more for loyalty than exceeding expectations does. For a chatbot, that is the most directly useful of the three metrics here.
How to Calculate CES:
Ask users to agree or disagree with a statement such as “the company made it easy for me to handle my issue”. Both 5-point and 7-point agreement scales are in common use; pick one and never change it, because scores across different scales are not comparable.
Formula:
CES = (Sum of scores) ÷ (Total responses)
An earlier version of this article stated that a “strong” CES falls between 5.5 and 7.0, and that below 5.0 signals a problem. That was unsourced and has been removed. What matters is the direction of travel against your own baseline and the difference between query types.
What to Monitor:
- Resolution Path Length: How many steps users take, and where they stall.
- Handoff Necessity: How often issues go to a human, and which types.
- User Input Clarity: How often users rephrase, which usually means the bot misread intent.
Tips to Improve CES:
- Simplify conversation flows: fewer steps for common tasks, direct routes to frequently requested information.
- Improve intent recognition: retrain on real failed conversations rather than imagined ones.
- Use smart defaults: pre-fill from context and history so the user types less.
When to Measure CES:
- Immediately after resolution
- After complex interactions
- After a chatbot-to-human handoff
- During a feature rollout
Best Practices:
- Keep the survey to one question where possible.
- Use the same wording and scale on every channel.
- Compare across interaction types, not against other companies’ published scores.
4. Total Conversations
Total Conversations shows how often the chatbot is used. On its own it is a vanity metric — volume rises when a product breaks as readily as when adoption grows. It is useful as a denominator for the rate metrics below, and as context for capacity planning.
5. Chat Duration
Chat Duration measures how long interactions run. Longer is not automatically worse and shorter is not automatically better, which is why a target figure is the wrong tool here.
Long durations can mean a clunky flow, unclear responses or unnecessarily deep paths. They can also mean the chatbot is handling something genuinely complicated that would otherwise have gone to a person. Track the median alongside the mean, because a few very long sessions distort the average.
The useful comparison is duration against resolution rate. Short chats with high resolution means efficiency. Short chats with low satisfaction means users are giving up.
Segment durations rather than setting a single target:
| Query Type | What to compare | What a change usually means |
|---|---|---|
| Basic FAQs | This month against your own baseline | Rising duration suggests answers are no longer matching the questions |
| Technical Support | Duration against resolution rate | Long and unresolved means the bot is out of its depth; hand off sooner |
| Sales Inquiries | Duration against conversion | Longer is fine if it converts; short and unconverted is a broken flow |
| Account Issues | Duration against handoff rate | Rising handoffs with rising duration means the bot is delaying, not helping |
An earlier version of this article gave a target duration in minutes for each of these categories. Those figures were not sourced from anywhere and have been removed.
6. Chat Exit Rate
Chat Exit Rate is the share of conversations users abandon before resolution.
Divide abandoned conversations by total interactions and multiply by 100. As arithmetic: 150 abandoned out of 1,000 gives a 15% exit rate. That is an illustration of the calculation, not a benchmark.
Read it together with response and resolution time. A spike in exits usually has a cause you can find in the transcripts.
7. Response Time
Response time is how quickly the chatbot replies. Delays cause drop-offs, and drop-offs show up in the exit rate.
Set a baseline from your own data and watch peak periods separately, since that is where infrastructure limits appear first.
8. Resolution Time
Resolution time covers the whole interaction, from first query to actual solution — not just the speed of the first reply.
Long resolution times usually point to complex issues, gaps in the knowledge base, or escalations that should not have been necessary.
To improve it:
- Update the knowledge base from real unresolved conversations.
- Analyse conversation flow on the slowest chats, looking for:
- Bottlenecks
- Misunderstood queries
- Answers that are technically correct but incomplete
- Set your own benchmarks by tracking averages for:
- Issue type (billing, technical support)
- Peak versus off-peak periods
- New versus returning customers
An earlier version of this article suggested aiming for a specific resolution time on simple queries. That figure had no source and has been removed. Derive your own target from your first month of data.
Flag interactions that exceed your target and read them. Speed matters less than whether the answer was right.
9. Self-Service Rate
The share of enquiries the chatbot resolves without a human.
Formula:
Self-service rate = (Conversations resolved by chatbot ÷ Total conversations) × 100
As arithmetic: 700 resolved out of 1,000 gives 70%. Again, an illustration of the formula rather than a target.
Be careful how “resolved” is defined in your analytics. Many platforms count any conversation that did not escalate as resolved, which quietly counts abandonment as success. Check that definition before you report the number to anyone.
10. Sales Success Rate
Formula:
Sales success rate = (Number of sales ÷ Total sales conversations) × 100
Worked example:
500 sales conversations with 75 purchases gives 15%. This is arithmetic, not a typical figure.
To track it usefully:
- Define what counts as a sale, in writing, before you measure.
- Tag sales conversations separately from support chats.
- Decide the attribution window — whether a purchase two days later counts, and why.
Break the rate down by product category, period and customer group. Attribution is the weak point of this metric: a chatbot often assists a sale that would have happened anyway, so treat the figure as directional.
11. Sales Amount
Formula:
Total Sales Amount = Sum of all purchases attributed to chatbot interactions
To track it:
- Tag chatbot-originated sales at source.
- Separate revenue from:
- Purchases completed in chat.
- Purchases started in chat and finished later.
- Upsells the chatbot suggested.
Practical tips:
- Use unique tracking codes for chatbot-driven purchases.
- State your attribution rule alongside the number, every time you report it.
- Break down by product, period and customer group.
- Compare average order value for chatbot and non-chatbot sales.
Related figures worth watching:
- Average order value from chatbot interactions.
- Revenue per chat session.
- Growth in chatbot-driven sales over time.
- Chatbot sales as a share of total revenue.
Account for seasonality and promotions before concluding the chatbot caused a change.
12. Cost Per Chat
Divide total chatbot cost — setup, subscription, maintenance, and the staff time spent training and correcting it — by conversations handled.
That last component is the one most calculations leave out, and it is often the largest. A chatbot that needs a person maintaining its knowledge base two days a week is not cheap, whatever the licence costs.
Cost per chat is only meaningful next to the self-service rate. Cheap chats that resolve nothing simply move the cost to your support queue.
Conclusion
Measuring chatbot ROI means holding efficiency metrics and satisfaction metrics side by side. Optimise response time alone and you will get fast, useless answers.
How to start:
- Set a baseline in your first month, before you change anything.
- Define what “resolved” means in your tooling, and check the platform agrees.
- Pick a small set of metrics you will actually review.
- Review monthly, and read transcripts as well as dashboards.
Where published figures exist, this article links to the source. Where they do not, it says so rather than supplying a number. If you find a chatbot benchmark quoted without a source, it is worth checking who first published it.
For tool comparisons and cost tracking, BizBot lists software in this category. Confirm pricing with the vendor before budgeting.
More on this topic
Browse all 21 articles on Customer Support, or jump straight to our buying guide: Best Customer Support Tools for Small Teams.
