← Back to blogs
Market note · 7 December 2024

Voice AI Agent Testing: Challenges and Opportunities

By Rahul Gupta, Alok Bishoyi, Shonik Agarwal
9 min read

There was a time when navigating automated voice menus meant pressing buttons and hoping to reach the right person. We have moved beyond that. Voice agents are evolving into true assistants, handling nuanced conversations and powering smart devices.

The evolution of voice AI agents

Advanced language models like GPT-4o have made these agents markedly more human-like, turning phones and devices into real gateways for managing appointments, customer service and beyond. Voice AI is already making inroads into banking, healthcare and hospitality — and we are just scratching the surface.

Voice 1.0 vs Voice 2.0

The move from Voice 1.0 to Voice 2.0 is more than technological advancement. It is a shift in how people interact with systems.

Voice 1.0 Voice 2.0
Flow Step-by-step; users pressed buttons following structured prompts Natural conversation combining ASR, LLMs, TTS and emerging speech-to-speech models
Paths Predefined and limited — predictable at the cost of flexibility Open-ended queries, dynamically adjusting to conversational nuance
Memory None between turns Context retained across interactions, allowing continuous improvement

Key challenges remain: background noise and interruptions disrupt performance, accent and dialect diversity is hard to handle accurately, agents still drift off-topic, and real-time processing latency degrades the experience. Moving from promising prototypes to products that scale reliably requires robust testing frameworks.

The stack behind voice AI

Speech-to-text (ASR)
Converts spoken words into text. Deepgram, Whisper, AssemblyAI.
Language processing (LLM)
Understands user intent and generates the response.
Text-to-speech (TTS)
Converts responses back into human-like speech. Eleven Labs, Azure.
Emotion engine
Adds tonal variation so conversations sound natural. Hume.
Streaming and telephony
Manages real-time interaction. LiveKit, Daily.

Multi-modal models are likely to consolidate some of these components, reducing complexity and cost while enabling more natural interactions. But each piece needs to be tested independently for robustness before integration into a cohesive system.

Full stack vs self-assembled solutions

Consideration Full stack (Retell AI, Hume, Vocode) Self-assembled
Complexity Infrastructure abstracted away; simpler to deploy More fine-tuned control, with added complexity
Flexibility Limited customization Full customization for differentiated experiences
Cost Higher per-call cost, competitive at scale Potentially more economical with technical proficiency
Control Opaque layers Visibility into every layer, easier troubleshooting

Testing: the current bottleneck

Building voice agents is one challenge; testing them effectively is often a bigger one. The existing landscape relies on labour-intensive methods — engineers making test calls and manually inspecting interactions, with feedback loops that take weeks. This is a significant bottleneck for scaling voice AI.

Rethinking testing for voice AI

Automation and simulation
Simulate thousands of conversations at scale in minutes to find edge cases and optimize performance.
Custom success metrics
Define and test what constitutes a successful voice interaction, aligned to real-world outcomes.
Rapid, detailed feedback
Actionable insight developers can iterate on, shortening development cycles.
Continuous monitoring
Real-time performance monitoring with alerts for deviation.

The key testing areas are functional testing (LLM outputs are unstructured, so slight phrasing variation changes responses — strong guardrails and extensive dynamic testing are required), performance evaluation, benchmarking with tools like VoiceBench, real-time interaction testing with frameworks like PipeCat, and quality assessment of synthesized speech.

Testing should be an enabler, not a bottleneck. Reducing manual intervention and shortening iteration cycles is key to unlocking voice AI.

A call to founders

Voice AI is advancing rapidly, but the current testing landscape is holding it back. The companies that solve these bottlenecks will set the standard for how voice interactions evolve across industries. If you are working on solutions that make voice AI testing seamless and scalable, we would love to collaborate — we bring not just capital but industry insight, technical expertise and strategic partnership.

Special thanks to Swapan (Haptik), Ashish (Convin) and Viswanath (Saarthi), whose insights helped shape this piece.

If you're a founder starting up, reach out at hi@dzero.vc.
← Back to blogs