Key Concepts:
- Voice Agents
- Evals (Evaluations)
- Self-Driving Car Development (Whimo)
- Simulation (Large-Scale)
- Probabilistic Evals
- Input/Output Evals
- Constant Eval Loops
- Regression Testing
- CI/CD (Continuous Integration/Continuous Deployment)
- Live Monitoring and Detection
- Level of Realism in Simulation
- Denoising
- LLM as a Judge
- Metric Calibration
- Benchmarking
- Task-Based Evals
- End-to-End Evals
1. The Problem: Voice Agents Are Not Everywhere Due to Lack of Trust
- Voice agents have the potential to automate critical workflows, but their adoption is hindered by a lack of trust.
- Enterprises overestimate what voice agents can do immediately while underestimating their potential in the near future.
- Scaling voice agents to production is difficult; they often get stuck in "POC hell" because enterprises are hesitant to deploy them for customer-facing issues.
- Two current approaches:
- Conservative/Deterministic: Forces the agent down a specific path (expensive IVR tree).
- Autonomous/Flexible: Allows the agent to handle new scenarios but is unpredictable and hard to scale.
- The speaker argues that reliability and autonomy are not mutually exclusive.
2. Learning from Self-Driving: The Whimo Example
- Whimo's success is attributed to large-scale simulation.
- Early manual evaluation methods (driving the car and noting issues) were not scalable.
- Specific tests for specific scenarios were brittle and expensive to maintain.
- The industry shifted to large-scale evaluation, focusing on the frequency of events across many simulations.
3. Similarities Between Self-Driving and Conversational Evals
- Both involve interacting with the real world and responding to the environment at each step.
- Simulations are crucial because responses vary based on inputs (e.g., "Hello, what's your name?" vs. "Hello, what's your email?").
- Static tests are expensive and break easily.
- Coverage is essential; simulate all possible scenarios to determine the probability of success.
- The non-determinism of LLMs is useful for simulating various responses.
4. Input/Output Evals vs. Probabilistic Evals
- Traditional LLM evals involve running inputs and evaluating outputs against a golden dataset.
- Conversational evals require reference-free evaluation, focusing on overall agent behavior rather than specific expected outputs.
- Examples of probabilistic metrics: How often does the agent resolve the user's inquiry? How often does it repeat itself? How often does it say things it shouldn't?
- Probabilistic evals enable scaling.
5. Constant Eval Loops for Scalability
- Constant eval loops are essential for making autonomous vehicles and voice agents scalable.
- Finding a bug, iterating on a fix, running regression tests, and implementing CI/CD workflows are crucial.
- The process includes:
- Reproducing the bug with evals.
- Fixing the issue.
- Running a larger regression set to ensure no new issues were introduced.
- Presubmit and postsubmit CI/CD workflows.
- Large-scale release testing (manual and automated).
- Live monitoring and detection.
- Manual evals are still important for human judgment calls.
6. Process for Voice Agent Development
- Start with simulated conversations and happy paths (e.g., booking an appointment).
- Run simulations, analyze data, identify failure points, and set up automated metrics.
- Iterate through this loop.
- Ship to production and run evals again.
- Flag issues for human review and feed the data back into simulations.
7. Level of Realism Needed in Simulation
- The required level of realism depends on what you're testing.
- Control variables and test for specific aspects.
- Hierarchy of realism:
- Workflows, tool calls, instruction following: Text-based tests are sufficient.
- Interruptions, latency, instructed pauses: Simple voices are sufficient.
- Accents, background noises, audio quality: Hyperrealistic voices are needed.
- Focus on the base-level components needed for development.
8. Denoising Technique
- Re-simulate failing scenarios multiple times to determine the probability of failure.
- Determine the acceptable level of reliability for different parts of the product (similar to cloud infrastructure's "nines of reliability").
9. Building an Eval Strategy
- Evals are a core part of product development, not just an engineering best practice.
- Defining metrics is equivalent to defining what the product should do well.
- Consider factors beyond latency, such as interruptions and adherence to instruction following.
- LLM as a judge can be flexible but noisy.
- Iterate on metrics and calibrate them with human feedback to increase confidence.
10. Steps for Approaching Evals
- Review public benchmarks.
- Benchmark with your own specific data (e.g., medical terms for a medical company).
- Run task-based evals (text-based or smaller modules).
- Run end-to-end evals at scale.
- Benchmark each part of the voice stack.
- Create an eval process with continuous monitoring, bug handling, and a hierarchy of test sets.
- Create dashboards and processes for continuous evaluation.
11. The Future of Voice AI
- Voice is the next platform shift, similar to web and mobile.
- Every enterprise will launch a voice experience, and user expectations will increase.
- The next generation of scalable voice AI will be built with integrated evals.
12. Notable Quotes:
- "I think the biggest problem to launching voice agents is trust."
- "...large scale simulation has been the huge unlock for self-driving and robotics..."
- "...eval are as important part of your process uh it's the key part of your product development and it's not just an engineering best practice this is actually like a core part of thinking through what does your product do"
- "voice unlocking all of these new really natural voice experiences it doesn't mean everything you should be doing via voice but there's really exciting potential there"
Conclusion:
The key takeaway is that building reliable and scalable voice agents requires a rigorous evaluation strategy inspired by the self-driving car industry. This involves embracing large-scale simulation, probabilistic evals, constant eval loops, and a focus on continuous improvement through data analysis and human feedback. By carefully considering the level of realism needed for different tests and establishing a well-defined eval process, enterprises can build trust in their voice agents and unlock their full potential.
AI summaries can miss context or contain errors. Check important details against the original video.





