Key Concepts
Finn Voice, AI agent, customer service, voice AI, knowledge-based agent, resolution rate, speech-to-text (STT), text-to-speech (TTS), Retrieval-Augmented Generation (RAG), telephony, conversation design, latency, escalation paths, context handoff, manual/automated evaluations, usage-based pricing, outcome-based pricing.
Finn Voice: An AI Agent for Phone Support
Introduction
The presentation focuses on Finn Voice, an AI-powered voice agent designed to handle frontline phone support, answering customer questions and escalating to human agents when necessary. The system was developed in approximately 100 days. The speaker discusses the development process and the potential of voice AI in customer service.
Intercom's Context
Intercom is a customer service platform and AI agent company, known for its messenger product. It has evolved into a comprehensive platform supporting channels like email, WhatsApp, and phone. Two years prior, Intercom launched Finn, an AI agent for text-based chat, which has seen significant growth with over 5,000 customers and resolution rates reaching 56% on average, and up to 70-80% for some customers. Finn includes tools for conversation analysis, agent training, and deployment. Finn Voice extends this system to phone calls.
Why Voice?
Voice remains a preferred channel for urgent or sensitive issues. Over 80% of support teams use phone support, and over one-third of customer service interactions occur via phone. Despite not being a legacy channel, phone support is costly, averaging $7-12 per call in the US with human agents. Voice AI agents can reduce costs by at least five times.
Benefits of Voice AI
- Availability: 24/7 support.
- No Wait Time: Instant availability, eliminating hold times.
- No IVR Menus: Natural speech interaction replaces traditional menu systems.
- Multilingual Support: AI agents can support 30-40+ languages.
- Cost Savings: Significant reduction in operational expenses.
- Scalability: Efficiently handles business growth and peak times.
Building Finn Voice: Seven Key Areas
The speaker outlines seven key areas that significantly impacted the development of Finn Voice:
- Use Case:
- Many voice AI startups focus on narrow problem spaces (e.g., scheduling appointments). Intercom opted for a flexible, knowledge-based agent capable of answering help article questions (e.g., pricing plans, return policies).
- This decision was based on evidence from Finn on chat, customer feedback, and analysis of call transcripts, which indicated that a large percentage of queries could be resolved with knowledge base content.
- The initial "wedge" use case was out-of-office hours support, replacing voicemail and allowing teams to test the technology without disrupting main workflows.
- Other use cases considered included authentication, information gathering (order ID, account ID), and smart routing, but these were not the primary focus initially.
- MVP Scope:
- The primary challenge was to ship a meaningful product quickly for testing with existing phone support customers.
- The focus was on three main experiences:
- Testing (Finn Voice Playground): A lightweight environment for simulating sessions and gathering feedback from customer service managers. Shipped within the first four weeks.
- Deployment: Allowing customer service managers to deploy the agent on phone lines with basic configuration options.
- Observability/Monitoring: Providing visibility into AI agent calls with transcripts, recordings, transcript summaries, and call outcomes for customer service agents.
- Text Stack:
- The core components involve a chain of STT, LM, and TTS.
- STT (Speech-to-Text): Converts speech to text.
- LM (Language Model): Generates the response.
- TTS (Text-to-Speech): Converts text back to audio.
- An alternative approach is voice-to-voice models, which process audio directly but offer less control over the output.
- Intercom started with a real-time API for rapid testing and evolved the stack while still using the real-time API as part of the core architecture.
- RAG (Retrieval-Augmented Generation): Critical for answering questions based on the knowledge base.
- Telephony: Integrating the agent with phone lines. Intercom had a head start due to its existing chat agent and native phone support product.
- Conversation Design:
- Voice is not simply chat with sound; the approach must be different.
- Key differences:
- Latency: Tolerance for delays is lower in voice. For simple queries, the response time was about 1 second. For complex queries (3-4 seconds), filler words ("let me look into this") were added to maintain conversation flow.
- Answer Length: Shorter, more concise answers are preferred in voice. Complex responses are broken into chunks, with the user prompted to confirm if they want to hear the next step.
- User Mindset: Some customers initially interacted with Finn Voice like an old-school IVR (single words). However, they adapted their behavior to use full sentences as they heard the agent doing so.
- Integration with Support Workflows:
- Feedback indicated that smooth integration with support team workflows was crucial.
- Key integration points:
- Escalation Paths: Configuring how calls are escalated to human support teams.
- Context Handoff: Generating transcript summaries after each AI agent call to provide context to the human agent.
- Evaluation:
- Manual and Automated Evals: Test conversations were run through every major code change, initially manually but with increasing automation.
- Internal Tooling: Internal web apps were built to review logs, transcripts, and recordings for troubleshooting.
- Resolution Rate: The north star metric, defined as either user confirmation of issue resolution or user disconnection after hearing at least one answer without calling back within 24 hours.
- LM as a Judge: Using another LM to analyze call transcripts to identify issues and opportunities for improvement.
- Pricing:
- Typical cost ranges from 3 to 20 cents per minute, depending on query complexity and provider.
- Pricing models:
- Usage-Based Pricing: Per minute or per call, predictable but doesn't capture agent quality.
- Outcome-Based Pricing: Charges only if the issue is resolved, aligning incentives but introducing risk for the provider. The market is expected to converge toward outcome-based pricing.
Final Thoughts
Finn Voice was built and deployed in approximately 100 days, with several enterprise customers using it on their main phone lines. Achieving the right performance (latency, resolution outcomes) is not just a model problem but also a product problem. It involves choosing the right use case, designing for phone conversations, building internal and external tools, integrating with support workflows, and building trust with decision-makers. The goal is to make the experience feel effortless, even with underlying complexity.
AI summaries can miss context or contain errors. Check important details against the original video.