Key Concepts
- Orpheus: Open-source, real-time voice model based on Llama 3B, outputs snack tokens.
- Snack: Open-source audio codec used by Orpheus.
- Low-Rank Adaptation (LoRA): Fine-tuning technique used for voice cloning.
- Head of Line Silence: Latency introduced by silence at the beginning of audio data in the Orpheus model.
- Batch Inference: Running multiple model generations concurrently on the same GPU.
- vLLM: Open-source library for serving language models, used for batch inference and dynamic quantization.
- Consistent Hash Ring: Load balancing strategy for distributing requests across multiple servers.
- WebRTC: Real-time communication protocol used for client connections.
Gabber's Experience Hosting Orpheus for Real-Time AI Personas
Neil, CTO of Gabber, discusses their experience hosting the Orpheus voice model for their real-time AI persona platform. Gabber focuses on consumer applications of AI personas, such as AI girlfriends, AI NPCs, AI therapists, AI personal trainers, and AI toys. They aim to make real-time synchronous AI experiences as ubiquitous as websites and apps.
The Need for In-House Voice Models
The high cost of existing voice platforms (upwards of $5 per hour) made them unsuitable for most consumer use cases. Gabber needed a solution that was close to free. They decided to host their own voice models to reduce costs and gain more control.
Orpheus: A Game Changer
Orpheus was the first open-source, real-time streaming voice model that met their needs. It's based on a Llama 3B model and trained on 100,000 hours of voice and text data. It outputs audio tokens called snack tokens, which are then decoded into 24 kHz audio. A key requirement is maintaining a throughput of 85-100 tokens per second to avoid gaps in the audio.
Voice Cloning with LoRA Fine-Tuning
Gabber uses LoRA fine-tuning for voice cloning. They found that one-shot cloning didn't work well with Orpheus due to its limited training data. LoRA allows them to create emotive and high-fidelity clones with relatively small amounts of training data. An example was shown of cloning Jack's voice using 10 minutes of data and a LoRA with rank 16 and alpha 32, trained for 5 epochs. While overfit, the result was still reasonably good.
Latency Optimization
Latency is critical for real-time voice applications. Four factors affect latency:
- Time to first token: The time it takes for the model to generate the first audio token.
- Tokens per second: The rate at which the model generates audio tokens.
- Network latency: The time it takes for data to travel across the network.
- Head of Line Silence: Silence at the beginning of the audio data.
Gabber found that head of line silence was a significant source of latency in the default Orpheus voice (e.g., Tara had 600ms of silence). Even filtering out the silence only saved 10% because the model was barely faster than real-time. They were able to fine-tune the silence away, reducing latency to around 100ms. This is important because it gives the LLM more time to generate tokens within the latency budget. They aim to generate the first audio packet within the "snooze period" (1-1.5 seconds) after the user stops talking.
Infrastructure with vLLM and Consistent Hashing
Gabber needed a robust and cost-effective infrastructure for hosting Orpheus. They chose vLLM because it supports batch inference with LoRAs, which allows them to run multiple generations and multiple LoRAs on the same GPU concurrently. The FP16 model was initially too slow on L40S GPUs, but vLLM's FP8 dynamic quantization solved this, bringing them to 105 tokens per second on non-fine-tuned voices and 95 tokens per second on LoRA voices with a batch size of 10.
For load balancing, they use a consistent hash ring to ensure that requests for the same session end up on the same GPU, which is important for caching LoRAs in memory and supporting streaming input and arbitrarily long generations. The hash ring distributes servers around a virtual ring, and requests are routed to the nearest server based on a hash of the session ID. This allows for easy scaling of popular clones by adding them to more servers.
System Architecture
The system architecture consists of:
- WebRTC backend: Terminates client connections.
- Websockets: Connects the WebRTC backend to the GPUs.
- Redis: Used to store the mapping between session IDs and GPUs.
When a session starts, the WebRTC backend connects to any GPU and queries Redis to determine the correct GPU for the session. It then proxies the request to the correct GPU using a TCP connection.
Conclusion
Gabber's experience demonstrates that it is possible to host voice models on GPUs and handle the infrastructure in a scrappy and cost-effective way. Open-source tools like Orpheus, vLLM, and LiveKit are enabling new and exciting use cases for real-time AI personas.
AI summaries can miss context or contain errors. Check important details against the original video.