Why CX Leaders Should Stop Measuring Voice AI by Deflection Alone

As voice AI moves into enterprise customer service, Parloa argues the real test is whether it can escalate cleanly and support better service outcomes

5
Sponsored Post
Contact Center & Omnichannel​Interview

Published: July 22, 2026

Nicole Willing

Voice AI is changing the customer service conversation, but enterprises still need to be careful about how they measure success. 

Contact centers have long used deflection as a simple way to understand the value of automation. If fewer calls reach human agents, the logic goes, the operation becomes more efficient as costs fall and queues shrink.  

That logic is understandable, but it is also incomplete. 

According to Dmitry Timofeev, Director of Product at Parloa, deflection became popular because customer service has often been viewed through a cost lens, an “attractive trap” for leaders. 

“It became a very common measure of success because historically customer service contact centers have been a cost center, and they’ve been treated as a cost center. Deflection directly maps to cost savings.” 

The danger in this approach is that the contact center optimizes for the wrong thing. A voice AI agent might answer instantly and reduce the number of calls that reach live agents. But none of that proves the customer’s issue has been resolved.  

Why Deflection Is Too Narrow for Voice AI 

In voice channels, customers often call because they have a specific, urgent, or complex issue. If the AI agent only blocks the path to a human agent, or repeats basic information without resolving the problem, the experience can deteriorate quickly. 

“The best way to ensure a hundred percent deflection is to simply hang up on a call without solving the customer problem. Then every single call will be deflected, but what you will have is a lot of angry customers who want to have nothing to do with your company anymore.” 

Deflection only tells part of the story. It does not show whether the customer received the right answer, understood the next step, avoided repetition, or reached a human agent with enough context for the issue to move forward. 

Timofeev argued that success needs to be defined from the customer’s perspective. 

“Every time we talk about customer support, what good looks like, you need to define it from the consumer perspective. Ideally, the person who is calling wants to get their problem solved painlessly and quickly, and that is more than just a deflected call.” 

This is why voice AI should be measured by resolution, Timofeev explained.  

“Companies who are incentivized on deflection will optimize for deflection at the expense of everything else, including at the expense of the customer experience, which is what you actually want to drive.” 

That point is becoming more important as voice AI adoption accelerates. 

The Real Test Is Resolution 

Customers are beginning to experience AI agents that connect quickly and manage more useful conversations than older IVR systems. That raises expectations across other brands and service environments. 

Timofeev argued that this presents an opportunity, rather than simply a cost to be reduced. 

“Every customer problem is an opportunity to invest into a customer relationship and make it better.” 

That shifts the role of voice AI. Rather than acting as a barrier between the customer and the contact center, voice AI should help move the customer towards a useful outcome. It should identify intent, manage the right journeys, complete suitable tasks, and escalate when the situation requires human support. 

Service teams should aim “to get customers instantly connected with an AI agent instead of hanging in the queue for a long time,” Timofeev explained. “And they want to make sure that the AI agent is able to solve as many problems as possible, starting from the simpler ones and highest volume ones, but then going further up and up and up into complexity.” 

Escalation Quality Matters 

This is where measurement needs to evolve. Deflection can still have a place, but it should not sit alone. Voice AI should also be judged by resolution rate, escalation quality, customer effort, sentiment, repetition, and the impact on human agents. 

AI agents will not solve every issue. Some conversations need a human agent because the case is sensitive, complex, unusual, or outside the AI agent’s configured scope. 

“AI agents generally can only resolve the issues that they are configured for,” Timofeev said. “So, there will always be room for bringing a human agent into the conversation.” 

A handover is not a failure in itself, but the problem comes when the handover is poor. 

A customer who has already explained their issue to an AI agent should not have to start again. The human agent should receive the context, the authentication status, what has already happened, and what the customer is trying to achieve. 

Without that, voice AI can make the human agent’s job harder. If customers have to repeatedly ask to speak to a representative, by the time they finally reach a person, the damage has already been done. 

Average Handling Time Can Mislead 

Voice AI also changes how contact centre leaders should interpret operational metrics. 

If AI handles simpler issues successfully, human agents may receive a higher proportion of complex cases. That can change traditional metrics such as average handling time, which can become misleading when voice AI is taking on simpler interactions and human agents spend more time on calls. 

“The average handling time might actually go up, while the total handling time, the overall handling time goes down,” Timofeev explained. “And the enterprises have the tools within their CCaaS systems to measure the handling time overall.” 

If leaders only track average handling time, they may conclude that performance has worsened. But a broader view of total human handling time, AI containment quality, and escalated call complexity can provide a more accurate picture of operational impact. 

Customer sentiment is another important signal. 

“Organizations should think of resolution rates, but also other things like sentiment and sentiment trajectory during the call,” Timofeev said.  

“If we made the initially angry customer neutral, this is an improvement by itself.” 

Customer effort also needs attention. Voice AI should reduce unnecessary repetition. If a customer has to repeat the same details to the AI agent and then again to the human agent, the experience has not been designed properly. “The rate of repetition is something that is important to measure,” Timofeev said. 

Building a Better Voice AI Measurement Framework 

The task for enterprise teams is to build a measurement framework that matches how voice AI works. Timofeev recommends using clear evaluation rules, rather than vague assessments of success. 

Some of those rules can be deterministic. For example, did the AI agent call the correct API, complete a workflow, authenticate the customer, or pass the right context to a human agent? 

Other assessments may need to evaluate conversation quality. That can include reviewing transcripts against defined criteria, including whether the customer’s issue was understood, whether the response was appropriate, and whether escalation happened at the right time. 

“These are non-deterministic systems that will never have 100% accuracy. But if you think about it, neither do humans.” 

Human quality checks still have a role. 

“You might still want to have humans taking a random sample of conversations and comparing the records,” he says. “How does your set of evaluations score a given transcript versus a human’s evaluation? You might want to iterate on your set of evaluations based on that kind of comparison with the human evaluation.” 

The Questions CX Leaders Should Ask 

Voice AI should be judged by whether it helps customers get closer to a resolution. 

That means looking beyond deflection and asking better questions. Was the issue solved? Was effort reduced? Was the customer sentiment improved? Was escalation handled cleanly? Did the human agent receive the right context? Did total operational effort fall, even if some individual calls became more complex? 

For enterprises, those questions offer a more balanced view of value. Deflection may still matter, but it should not be the headline measurement. 

As Timofeev put it, “this is not just a cost-cutting initiative, but an investment into customer retention.”

Agent AssistAI Voice AssistantsCall & Contact Center SoftwareConversational AI
Featured

Share This Post