Build a Real-Time AI Sales Agent - Sarah Chieng & Zhenwei Gao, Cerebras
By AI Engineer
Key Concepts
- Cerebras Hardware (WSE-3): A wafer-scale engine with 4 trillion transistors and 900,000 cores, designed to overcome memory bandwidth limitations in AI inference.
- Speculative Decoding: A technique combining a smaller, faster "draft" model with a larger, more accurate model to accelerate inference.
- LiveKit: A real-time infrastructure platform using WebRTC for low-latency voice data transmission, essential for building voice agents.
- Voice Agent: A stateful, intelligent system capable of natural language understanding, real-time response generation, and maintaining conversational context.
- STT (Speech-to-Text): Converts spoken language into text.
- TTS (Text-to-Speech): Converts text into spoken language.
- VAD (Voice Activity Detection): Detects when a speaker is actively speaking.
- LLM (Large Language Model): The core "brain" of the agent, responsible for understanding and responding to user input.
- Context Loading: Providing the LLM with specific information about the business (product details, pricing, objection handling) to improve accuracy and relevance.
- Multi-Agent System: Utilizing multiple specialized agents (greeting, sales, technical, pricing) to handle different aspects of a conversation.
- Tool Calling: Enabling the agent to access and utilize external tools or APIs to fulfill user requests.
Hardware Innovations at Cerebras
The presentation highlighted Cerebras’ approach to overcoming limitations in traditional AI hardware, specifically focusing on memory bandwidth bottlenecks in GPUs. NVIDIA’s H100 GPU architecture relies on cores accessing weights, activations, and KV cache from off-chip memory, creating a significant bottleneck. Cerebras’ WSE-3 (Wafer Scale Engine) addresses this by providing each of its 900,000 cores with dedicated on-chip SRAM memory. This direct access dramatically reduces latency and accelerates inference speeds, achieving 20x-70x faster performance compared to NVIDIA GPUs. Sarah Chang demonstrated this speed advantage with a live Llama model prompt, showcasing rapid generation of lengthy responses. Beyond hardware, Cerebras also employs software optimizations like speculative decoding, combining a smaller, faster model with a larger, more accurate one to further enhance performance. The smaller model generates tokens quickly, while the larger model verifies their accuracy, leveraging the strengths of both.
Building a Voice Sales Agent: A Step-by-Step Process
The workshop detailed a four-step process for building a voice sales agent using Cerebras technology and LiveKit’s agent SDK:
Step 1: Setup & API Keys: Install necessary packages (LiveKit agents, Cartisia, Cilero, OpenAI compatibility) using the provided starter code. This establishes the foundational infrastructure.
Step 2: Context Loading: "Train" the agent by providing it with relevant business information. This involves loading product descriptions, pricing details, and pre-written responses to common objections into a structured format accessible to the LLM. The goal is to minimize hallucinations and ensure accurate, on-message responses. The example code demonstrates how to organize this information.
Step 3: Agent Creation: Wire together the various components (LLM, TTS, STT, VAD) using the LiveKit agent SDK. The SalesAgent class encapsulates this logic, loading the context, defining rules for voice interaction (avoiding bullet points, prioritizing context), and initializing the agent with the necessary configurations. An onenter method triggers the conversation when a user connects.
Step 4: Launching & Expanding: Start the agent by connecting it to a virtual room, creating an instance of the SalesAgent, and initiating a session. The presentation then discussed expanding the agent’s capabilities through:
- Multi-Agent Systems: Implementing specialized agents (greeting, sales, technical, pricing) and a handoff mechanism to route users to the appropriate expert.
- Tool Calling: Enabling the agent to access external tools and APIs to fulfill user requests (e.g., checking inventory, processing orders).
Voice Agent Architecture & Functionality
A voice agent operates through a three-phase process:
- Listening Phase (STT & VAD): Speech is detected and converted to text using Speech-to-Text (STT) technology. Voice Activity Detection (VAD) helps prevent interruptions by accurately identifying when a speaker has finished talking, utilizing a smaller model to predict utterance completion.
- Thinking Phase (LLM & Context): The transcribed text is sent to a Large Language Model (LLM) for understanding and response generation. The LLM leverages the loaded context (business information) to provide relevant and accurate answers.
- Speaking Phase (TTS): The LLM’s response is converted back into speech using Text-to-Speech (TTS) technology and streamed back to the user in real-time.
LiveKit orchestrates these components, managing audio streams, maintaining conversational context, and coordinating the AI services. The use of WebRTC ensures low-latency communication.
Key Arguments & Perspectives
- Hardware Matters: Cerebras argues that hardware innovation is crucial for advancing AI, particularly in inference. Their WSE-3 chip addresses the memory bandwidth bottleneck that limits GPU performance.
- Voice is the Future of Interaction: The presenters believe voice agents will become increasingly prevalent in customer service, sales, and other applications, offering a more natural and efficient communication experience.
- Context is King: Providing LLMs with specific, relevant context is essential for generating accurate and helpful responses, especially in specialized domains like sales.
- Multi-Agent Systems Enhance Capabilities: Specialized agents, working together with a handoff mechanism, can provide a more comprehensive and effective customer experience.
Notable Quotes
- “Cerebras chips do not have memory bandwidth issues.” – Sarah Chang, highlighting the core advantage of their hardware.
- “LLMs are only as good as their training set.” – Genway, emphasizing the importance of context loading.
- “The best way to really have these customer interactions is through real conversations, which is why voice agents are so powerful.” – Genway, articulating the value proposition of voice agents.
Data & Statistics
- Cerebras WSE-3: 4 trillion transistors, 900,000 cores.
- Performance Comparison: Cerebras claims 20x-70x faster inference speeds compared to NVIDIA GPUs.
- Latency: LiveKit utilizes WebRTC to achieve less than 100 milliseconds of latency.
- Llama 3.3: Used as the LLM in the workshop, benchmarked by Artificial Analysis as competitive with other models.
Logical Connections
The presentation flowed logically from a discussion of hardware innovation (Cerebras) to the software infrastructure (LiveKit) and finally to the practical application of building a voice sales agent. The hardware segment provided the foundation for understanding the performance benefits, while the software segment explained how to orchestrate the various AI components. The step-by-step workshop then demonstrated how to leverage these technologies to create a functional agent. The expansion section (multi-agent systems, tool calling) built upon the core agent functionality, suggesting future development possibilities.
Conclusion
The workshop provided a comprehensive overview of building voice agents, emphasizing the importance of both hardware and software innovation. Cerebras’ WSE-3 chip offers significant performance advantages, while LiveKit’s agent SDK simplifies the orchestration of complex AI components. By focusing on context loading, multi-agent systems, and tool calling, developers can create powerful and versatile voice agents capable of transforming customer interactions. The provided code notebook and API credits empower attendees to immediately begin building and deploying their own solutions.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Stanford CS153 Frontier Systems | Building the Frontier Ecosystem
Stanford Online

'Things are going to be okay, in Canada and the U.S.': Thorne
BNN Bloomberg

I'M OUT: The $11 Trillion AI Bubble is Breaking!
Steven Van Metre

South Korea bets big on AI with nearly a trillion dollars of investment • FRANCE 24 English
FRANCE 24 English

The Bubble is Bursting... (Emergency Update)
Bravos Research

The AI Bubble Just Ended - Without Popping
Heresy Financial

AI Market Volatility, Europe Heat Wave, Venezuela Quakes Damage | Bloomberg This Weekend: June 27
Bloomberg Television