Building voice agents with OpenAI — Dominik Kundel, OpenAI

AI EngineerAbout 7 min readJun 30, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Voice Agents: Systems that accomplish tasks independently on behalf of users through voice interaction.
  • Agents SDK (TypeScript): An SDK providing abstractions for building agents based on OpenAI's best practices, including features like handoffs, guard rails, streaming, tool calling, and tracing.
  • Real-time Agents: Specialized agent configurations for voice interactions, handling audio input/output, guard rails, and life cycle management.
  • Ephemeral Key (Client Secret): A short-lived API key generated by the server and passed to the client for secure interaction with the real-time API in browser-based applications.
  • Speech-to-Text (STT) and Text-to-Speech (TTS) Chained Approach: An architecture where audio is converted to text, processed by a text-based agent, and then converted back to audio.
  • Speech-to-Speech Approach: An architecture where a model trained on audio directly interacts with the conversation, bypassing transcription.
  • Handoffs: Transferring control between different agents with specific skills or roles during a conversation.
  • Guard Rails: Mechanisms to ensure agents adhere to policies and avoid inappropriate behavior.
  • Delegation: Using tool calls to delegate complex tasks to more specialized or powerful models.
  • Traces API/Dashboard: A tool for debugging and monitoring agent behavior, including audio, transcripts, and tool calls.
  • Zod: A library for defining schemas to validate the arguments that the model is going to receive.

1. Voice Agents and the Agents SDK

  • Definition of Agents: Systems that accomplish tasks independently on behalf of users, combining a model with instructions, access to tools, and a runtime for lifecycle management.
  • OpenAI Agents SDK for TypeScript: Launched to provide abstractions for building agents based on best practices. It includes features like handoffs, guard rails, streaming input/output, tool calling, built-in tracing, human-in-the-loop support, and native voice agent support.
  • Native Voice Agent Support: Allows building voice agents with the same primitives as the Agents SDK, handling handoffs, output guard rails, context management, and built-in tracer support. It also includes native interruption support and WebRTC/WebSocket support.

2. Why Voice Agents?

  • Accessibility: Voice agents make technology more accessible to a wider range of users.
  • Information Density: Voice communication is more information-dense than text due to tone, voice, and emotions.
  • API to the Real World: Voice agents can interact with businesses or systems that lack traditional APIs.

3. Voice Agent Architectures

3.1. Chained Approach (STT -> Text-based Agent -> TTS)

  • Process: Audio is converted to text using STT, processed by a text-based agent, and then converted back to audio using TTS.
  • Strengths: Easier to get started with existing text-based agents, full access to any model, and more control/visibility of model actions.
  • Challenges: Turn detection, added latency, and loss of audio context.

3.2. Speech-to-Speech Approach

  • Process: A model trained on audio directly interacts with the conversation, bypassing transcription.
  • Strengths: Lower latency and more contextual understanding of audio.
  • Challenges: Reusing existing text-based capabilities and dealing with complex states/decision-making.
  • Solution: Delegation using tools, where a frontline agent interacts with the user and uses tool calls to interact with smarter reasoning models.

3.3. Demo of Real-time Agent

  • A real-time agent built with the Agents SDK was demonstrated, showcasing its ability to handle interruptions and call tools.
  • The agent was able to process a refund request by calling a more advanced tool (GPT-4 Mini) and provide a response.
  • The traces UI was used to show the audio, tool calls, and interactions between the front-end and back-end agents.

4. Best Practices for Building Voice Agents

  • Start with a Small and Clear Goal: Focus on solving a specific problem with a limited number of tools.
  • Build Evals and Guard Rails Early On: Use the traces API/dashboard or custom dashboards to monitor performance and ensure safety.
  • Prompting for Tone and Voice: Use generative models to prompt for specific tones, voices, emotions, and personalities.
  • Conversation States: Use JSON structures to define conversation flows and guide the model through specific steps.

5. Building a Voice Agent with the Agents SDK

5.1. Setting up the Project

  • Clone the provided GitHub repository and install the dependencies.
  • The repository includes a boilerplate Next.js app and an empty package.json file.

5.2. Creating a Basic Agent

  • Import the agent class from the @openai/agents package.
  • Define an agent with instructions, a name, and a run function.
  • Log the final output of the agent.
  • Example:
    import { agent, run } from '@openai/agents';
    
    const myAgent = new agent({
      name: 'My Agent',
      instructions: 'You are a helpful assistant.',
    });
    
    const result = await run(myAgent, 'Hello, how are you?');
    console.log(result.finalOutput);
    

5.3. Adding Tools

  • Import the tool class and define a tool with a Zod schema for argument validation.
  • Give the tool to the agent.
  • Example:
    import { tool } from '@openai/agents';
    import { z } from 'zod';
    
    const getWeatherTool = new tool({
      name: 'getWeather',
      description: 'Gets the weather for a given location.',
      args: z.object({
        location: z.string().describe('The location to get the weather for.'),
      }),
      async run({ location }) {
        // Code to get the weather
        return `The weather in ${location} is sunny.`;
      },
    });
    
    myAgent.tools = [getWeatherTool];
    

5.4. Building a Real-time Agent

  • Import the realTimeAgent class from the @openai/agents/realtime package.
  • Define a real-time agent with a name and instructions.
  • Use an ephemeral key (client secret) for browser-based applications.
  • Create a real-time session with the agent and connect to it using the API key.
  • Example:
    import { realTimeAgent } from '@openai/agents/realtime';
    import { createRealTimeSession } from '@openai/agents/realtime/server';
    
    const myRealTimeAgent = new realTimeAgent({
      name: 'My Real-time Agent',
      instructions: 'You are a helpful assistant.',
    });
    
    const session = await createRealTimeSession({
      agent: myRealTimeAgent,
      model: 'gpt-4-1106-preview',
    });
    
    await session.connect({ apiKey: clientSecret });
    

5.5. Displaying the Transcript

  • Listen to the historyUpdated event to get the conversation history.
  • Filter the history to show only messages.
  • Display the messages in a list.

5.6. Handling Interruptions

  • The real-time session handles interruption events automatically.
  • The transcript is updated to reflect the interruption.

5.7. Human Approval

  • Use the needsApproval property to require human approval before executing a tool.
  • Listen to the toolApprovalRequested event to get the approval request.
  • Approve or reject the tool call.

5.8. Handoffs

  • Create a specialized tool call that resets the configuration of the agent in the session.
  • Update the system instructions and tools.

5.9. Delegation

  • Create a separate agent on the back end to handle complex tasks.
  • Use a tool call to delegate the task to the back-end agent.

5.10. Guard Rails

  • Use guard rails to protect the input and output of the agent.
  • Define guard rails to check for policy violations.
  • Interrupt the agent if a violation is detected.

6. Additional Points

  • The real-time API holds the source of truth for the conversation session.
  • The client receives a copy of the events.
  • The updateHistory event can be used to update the session context.
  • The API provides detailed information about token usage.
  • The speed parameter can be changed mid-session.
  • The API will throw an error if a parameter cannot be changed.
  • The real-time API supports switching languages.
  • There is no speaker detection in the model.
  • Custom voices are not currently supported.
  • There is no prompt caching control with real time.
  • There is no wake word detection built in.

7. Conclusion

The presentation provides a comprehensive overview of voice agents and the OpenAI Agents SDK for TypeScript. It covers the key concepts, architectures, best practices, and implementation details for building voice agents. The Agents SDK simplifies the development process by providing abstractions for common tasks such as handoffs, guard rails, and tool calling. The presentation also highlights the importance of evaluating and monitoring voice agents to ensure they meet performance and safety requirements.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.