Build Your First Voice AI Agent in 20 Minutes with LiveKit (Open Source)

By Cole Medin

Share:

Key Concepts

  • Voice AI Platforms: Tools like Vappy, Synthflow, and Bland.ai for building voice agents, often criticized for their "black box" nature and limitations.
  • LiveKit: An open-source Python framework for building highly customizable, self-hostable, and scalable voice agents.
  • Voice Pipeline: The sequential process in a voice agent involving Speech-to-Text (STT), a Large Language Model (LLM), and Text-to-Speech (TTS).
  • Agent Class: A Python class (e.g., Assistant) that defines the structure and behavior of a LiveKit voice agent.
  • Agent Session: An instance of an agent's interaction, managing parameters and the conversation flow.
  • @function_tool Decorator: A Python decorator used in LiveKit to designate a function within an agent class as a callable tool for the LLM.
  • Docstring: A string literal used to document Python code, which LiveKit agents leverage to understand when and how to use a tool, including its arguments.
  • MCP (Multi-Agent Communication Protocol) Server: A server that allows LiveKit agents to integrate with external services and APIs, often via protocols like Streamable HTTP.
  • LiveKit CLI: The Command Line Interface for interacting with LiveKit, used for authentication, environment setup, and agent deployment.
  • Self-hosting: The ability to run LiveKit infrastructure on one's own servers for complete control.
  • RAG (Retrieval Augmented Generation): An advanced AI technique where an LLM retrieves information from an external knowledge base to enhance its responses.

Critique of Existing Voice AI Platforms

The video begins by highlighting the trade-offs of popular Voice AI platforms such as Vappy, Synthflow, and Bland.ai. While these platforms are easy to use and powerful for getting started, they present significant limitations for businesses requiring deep customization and control. Key issues include:

  • Lack of Infrastructure Control: Users do not run their own infrastructure, leading to dependency on the platform's setup.
  • Slow Tool Calls: Integrations and external function calls can be sluggish.
  • Premium Per-Minute Rates: Cost can escalate rapidly due to per-minute pricing models.
  • Limited Customization: It is challenging to truly customize agents beyond predefined options, making them a "big black box."

The presenter notes that businesses have specifically switched from platforms like Vappy to custom solutions due to these problems, indicating a need for more flexible alternatives.


Introducing LiveKit: An Open-Source Solution

LiveKit is presented as a superior alternative, an open-source Python framework designed for building voice agents with code. Its core advantages address the shortcomings of other platforms:

  • Full Customization and Control: Developers have complete control over conversation logic and agent behavior.
  • Direct Integrations: Seamless integration with custom tools and MCP servers.
  • Deployment Flexibility: Agents can be self-hosted for complete infrastructure control or deployed to the LiveKit cloud.
  • Performance and Scalability: LiveKit is described as fast, reliable, and highly scalable.
  • Ease of Use: Despite being code-based, LiveKit is surprisingly easy to get started with, allowing customization of tools and models for Speech-to-Text (STT) and Text-to-Speech (TTS).

Building a Basic LiveKit Agent Locally

The video demonstrates building a basic LiveKit agent in 52 lines of Python code, emphasizing its simplicity.

Step-by-Step Process:

  1. Import Dependencies: Import necessary modules from LiveKit, including providers for the voice pipeline.
  2. Load Environment Variables: Load API keys (e.g., OpenAI API key) and other configurations.
  3. Create an Agent Class: Define a Python class (e.g., Assistant) that inherits from LiveKit's Agent class. All agent logic resides within this class.
  4. Initialize the Agent (__init__ function):
    • Specify the system prompt for the voice agent, defining its initial persona and instructions.
  5. Define the Entry Point (entrypoint function):
    • This function is called when the agent is first invoked.
    • Create an agent session, where parameters for the voice pipeline are defined. The voice pipeline typically involves:
      • Speech-to-Text (STT) model: Transcribes user speech into text (e.g., Deepgram).
      • Large Language Model (LLM): Processes the text and generates a response (e.g., OpenAI's GPT models).
      • Text-to-Speech (TTS) model: Converts the LLM's text response back into spoken language.
    • LiveKit supports various providers for STT (e.g., Cartesia), LLM (e.g., Anthropic), and TTS, offering flexibility.
    • Start the session by creating a room, which maintains the conversation history between the user and the agent.
    • Generate an initial greeting using the agent's reply generation capability, allowing the agent to speak first.
  6. Run the Agent: Call CLI.run_app with the defined entry point to start the agent.

Local Testing:

The agent is tested locally using the console command via the LiveKit CLI. The agent successfully provides a greeting and answers a question about LiveKit's coolness ("scalable, real-time, open-source video").


Enhancing Agents with Tools

The video demonstrates how to easily add custom tools to a LiveKit agent, enabling it to perform specific actions or access external information.

Step-by-Step Process:

  1. Create a Python Function: Define a standard Python function within the agent class.
  2. Add @function_tool Decorator: Apply the @function_tool decorator from LiveKit to the function. This tells the LiveKit agent that this function is a capability it can use.
  3. Use Docstrings for LLM Instructions: The function's docstring is crucial. It informs the LLM when and how to use the tool, including specifying arguments. This works similarly to how tools are defined in frameworks like Pydantic AI or Crew AI.

Examples:

  • get_current_date_time Tool: A simple tool to retrieve the current date and time, as LLMs typically lack real-time information due to their training cutoff.
    • Demonstration: The agent successfully provides the current time (e.g., "4:21 p.m. on October 3rd, 2025") after being asked.
  • Airbnb Assistant with Mock Data:
    • search_airbnbs Tool: Takes a city as a parameter and returns mock Airbnb listings. The LLM decides the city parameter based on the conversation.
    • book_airbnb Tool: Takes name, check_in_date, and check_out_date as parameters. If any arguments are missing, the agent clarifies with the user.
    • Demonstration: The agent successfully searches for Airbnbs in San Francisco, recommends a "Cozy Downtown Loft," and then proceeds to "book" it after clarifying user details.

Integrating with MCP Servers for Real-World APIs

LiveKit agents can integrate with MCP (Multi-Agent Communication Protocol) servers to connect to real-world APIs.

Example: Real Airbnb API Integration

  1. MCP Server Setup: The presenter uses a Docker MCP gateway to run an Airbnb MCP server locally, exposing it via the Streamable HTTP protocol on port 8089. This setup will be detailed in a future video on the Docker MCP catalog.
  2. LiveKit Agent Configuration:
    • In the agent session definition, a list of mcp_servers is added.
    • For a local, unauthenticated server, only the URL (e.g., http://localhost:8089) is required.
  3. Demonstration: The agent is invoked and asked to search for Airbnbs in Minneapolis, Minnesota. It successfully uses the real Airbnb API to return a listing (e.g., "studio in historic Brownstone, downtown MLS, rating of 4.79"). The presenter clarifies that this integration performs real searches but does not actually book Airbnbs.
  4. Custom Logic: The video briefly mentions that LiveKit allows for custom logic to be built around conversation events, such as when a user starts or stops speaking, further extending customization beyond what other platforms offer.

Cloud Deployment and Browser Interaction

The final section covers deploying a LiveKit agent to the cloud and interacting with it via a web browser, as well as mentioning phone integrations.

Step-by-Step Cloud Deployment:

  1. Sign Up for LiveKit Cloud: Create an account on the LiveKit cloud platform.
  2. Install LiveKit CLI: Install the LiveKit Command Line Interface (CLI) using platform-specific commands (e.g., winget for Windows).
  3. Authenticate with LiveKit Cloud: Run lk cloud auth to authenticate the CLI with the LiveKit cloud account.
  4. Set Up Environment Variables: Use lk app env to enter environment variables (e.g., OpenAI API key, Deepgram API key, LLM choice). The CLI can pre-fill values from an .env.example file.
  5. Start the Agent: Run lk agent start. This command sets up necessary configurations under the hood. The process can be exited immediately after it loads.
  6. Create and Deploy the Agent: Run lk agent create.
    • Select the organization and the .env file for secrets.
    • Choose the agent script to deploy (e.g., the basic agent with mock tool calls, as MCP servers cannot be used remotely in this free tier example).
    • The CLI then creates a Docker file and deploys the agent to the LiveKit cloud.

Browser Interaction:

  • Once deployed, the agent is visible in the LiveKit cloud dashboard.
  • Users can interact with the agent directly in the LiveKit playground in their browser.
  • Demonstration: The deployed agent successfully responds to a request to find the top Airbnb in San Francisco, providing the same mock data response as the local version.

Phone Integration:

The presenter mentions that LiveKit supports telephone integrations, allowing agents to be connected to a phone number, which is often the end goal for many voice agents. This is a potential topic for a future video.


Conclusion and Future Possibilities

The video concludes by reiterating the ease, power, and customization capabilities of LiveKit for building voice agents using Python code. The presenter expresses enthusiasm for LiveKit and encourages viewers to explore its potential.

Future topics for LiveKit content could include:

  • Multi-agent workflows.
  • Advanced tool integrations.
  • A dedicated video on phone integrations.
  • A workshop on advanced LiveKit agents with RAG (Retrieval Augmented Generation) implementation (already covered in the Dynamis community).

The presenter encourages likes and subscriptions for more content on voice agents and AI agents.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video