Resolution rate tells you how much work an AI support agent handled, not whether it got the outcome right. Here are the metrics, and the escalation behavior, that actually build trust.
Every AI support vendor is selling you the same number: resolution rate.
"We resolve 60% of tickets." "Our agent handles 80% of volume." "Automated resolution across all channels."
It is the default metric for evaluating AI agents in customer support. It is also one of the least useful ones.
Resolution rate tells you how much work the agent handled. It does not tell you how much trust the agent earned, how many tickets it closed prematurely, how many customers had to reopen a conversation the agent said was finished, or how much hidden cleanup work it created for your human team.
A high resolution rate can mean the agent is performing well. It can also mean the agent is calling things "resolved" that are not actually resolved.
The problem with resolution as a north star
Resolution, as most AI vendors define it, means: the agent sent a response and the customer did not reply.
That is not the same as solving the customer's problem.
A customer might not reply because the answer was genuinely helpful. They also might not reply because the answer was confusing. Or because they gave up. Or because they went to a different channel. Or because the agent confidently stated something incorrect and the customer took it at face value.
When resolution rate is the primary metric, the incentive is to close tickets, not to get the right outcome. That distinction matters more than most teams realize, especially for complex or high-stakes support.
The harder question is not "how many tickets did the AI handle?" The harder question is: "when the AI handled a ticket, did it actually get the outcome right?"
What actually builds trust
Customers do not need the AI agent to solve everything. They need the agent to know when it should not try.
Support teams do not need the agent to eliminate every human touch. They need the agent to hand off the right work, with the right context, at the right moment.
The thing that builds trust is not resolution. It is escalation.
Proper escalation tells the customer:
- The agent knows when not to guess
- The agent will not pretend to have authority it does not have
- The company still has a human safety net
- The handoff will not strand the customer or make them start over
Proper escalation tells the support team:
- The agent understands its boundaries
- The agent is not silently creating messes
- The agent can be given more responsibility safely
- Human work is being routed intentionally, not dumped randomly
Bad escalation behavior is often worse than no automation. It creates hidden cleanup work. The human team spends time undoing things the agent should not have done, re-reading conversations the agent mishandled, and rebuilding customer trust the agent damaged. None of that shows up in the resolution rate.
Resolution is not one thing
One of the biggest mistakes in AI support evaluation is treating resolution as binary. Either the agent resolved it or it did not.
In practice, there are multiple valid outcomes for any ticket:
- Resolved: the agent fully completed the customer's request and no further action is needed.
- Deflected / answered: the agent provided a useful answer or next step, but the conversation may continue.
- Pending: the agent asked for more information or gave instructions and is waiting on the customer.
- Escalated: the agent handed the ticket to a human for review, approval, or action.
- Ended silently: the agent recognized the ticket was outside its scope and did not respond.
A strong AI agent should be able to distinguish between all of these outcomes and take the right action for each one.
The dangerous pattern is when the agent treats every answer as a resolution. It sends a helpful reply and marks the ticket solved, even when the customer still has an open question, when a human review is needed, when a promised follow-up has not happened, or when the agent was only confident enough to explain a likely cause but not confident enough to confirm a fix.
That distinction matters. The agent might be confident enough to explain what is probably happening, but not confident enough to issue a refund. It might be confident enough to ask for a log file, but not confident enough to close the ticket. It might be confident enough to say "this looks like a configuration issue," but not confident enough to say "this is resolved."
Confidence should not be a vibe. It should be tied to the action the agent is about to take.
The metrics that actually matter
Instead of leading with resolution rate, teams should track the metrics that reveal whether the agent is behaving correctly at the boundaries.
False resolution rate: how often the agent marks something resolved when it should not have. This is the single most corrosive failure mode. It erodes customer trust, creates reopens, and makes the human team skeptical of the AI.
Missed escalation rate: how often the agent fails to escalate when it should have. This includes cases where a human review was needed, a policy judgment was required, or the agent lacked the authority or information to complete the request.
Over-escalation rate: how often the agent escalates when it could have safely handled the ticket. This matters for efficiency, but it is a much safer failure mode than missed escalation.
Reopen rate after AI resolution: how often customers come back after the agent said the issue was finished. A high reopen rate is the clearest signal that resolution rate is inflated.
Human override rate: how often a human agent changes the outcome after the AI acted. This reveals cases where the AI took the wrong action and a human had to clean it up.
Customer reply after AI resolution: how often a customer replies to a ticket the AI marked as resolved. Not every reply means the resolution was wrong, but a pattern of post-resolution replies is a red flag.
Correct outcome classification: across all tickets, how often does the agent correctly distinguish between resolution, deflection, pending, and escalation? This is the foundational metric. If the agent cannot reliably classify the right outcome, the resolution number is meaningless.
What good escalation actually looks like
Good escalation is not just "the agent gave up." It is a deliberate, well-structured handoff.
The agent recognizes the boundary. It does not guess. It does not overstate its authority. The customer gets clear language about what happens next. The human team receives useful context, not a bare ticket with no explanation. The ticket lands in the right queue. Future AI replies are blocked so the agent does not keep jumping in after handoff.
The boundary instructions should be specific, not vague:
- Escalate when a human review, approval, refund, discount, account change, or manual action is required.
- Escalate when the customer asks for something outside documented policy.
- Escalate when a tool call fails and the customer still needs action.
- Escalate when the agent cannot determine whether the customer is eligible for something.
- Do not mark a ticket resolved just because a helpful answer was provided.
- If the response promises that a human will check, review, approve, or follow up, treat the outcome as escalation, not resolution.
That last point is especially important. If the agent's response says "our team will look into this," the ticket is not resolved. It has been escalated with a promise attached. Calling it a resolution inflates the number and breaks the workflow.
The improvement signal resolution rate hides
There is a deeper problem with leading on resolution rate. It hides the operational signal that would actually make the support system better.
A strong AI agent should not just work tickets. It should reveal where the support system itself is underbuilt.
Every escalation, every failure, every uncertain outcome contains information:
- Context gap: the agent lacked the right customer, account, or usage data to complete the request.
- Tooling gap: the agent needed access to a system, action, or integration it did not have.
- Permission gap: the agent could explain the situation but was not authorized to take the required action.
- Policy gap: the rules were too ambiguous to automate safely and required human judgment.
- Knowledge gap: documentation was missing, outdated, or insufficient.
- Workflow gap: the process required manual steps, unclear ownership, or a handoff path that was not defined.
- Product gap: repeated product confusion, edge cases, or UX issues were driving avoidable support volume.
Resolution rate collapses all of this into one number. It tells you how much work the agent handled, but not what is blocking the next level of performance.
The better question is not "what is our resolution rate?" The better question is: "what should we improve this week so the agent, the human team, and the customer experience all get better?"
The trust test
The real evaluation of an AI support agent is not "how many tickets can it close?"
It is: "can the agent stop before guessing, choose the right kind of stop, and give the human team enough context to continue without starting over?"
An agent that resolves 50% of tickets correctly and escalates the other 50% cleanly is more valuable than an agent that claims 80% resolution but closes 15% of tickets prematurely, creates reopens, and makes the human team spend time on cleanup instead of real support work.
Resolution rate is easy to inflate. Trust is not.
The best AI support agents do not chase the highest resolution number. They build the kind of operational trust that lets the support team confidently expand what the agent is allowed to do next. That is how resolution rate grows sustainably, not by counting more aggressively, but by genuinely becoming more capable.
The metric that matters is not how much the agent handled. It is whether the team trusts it enough to give it more.
