Diffusion LLMs are here...
By David Ondrej
Key Concepts
- Diffusion LLM: A model architecture that generates text through parallel refinement rather than sequential token generation.
- Parallel Refinement: The process of generating multiple tokens simultaneously and iteratively improving them, as opposed to the standard autoregressive (one-by-one) approach.
- Latency Compounding: The cumulative delay experienced in AI agents when multiple sequential LLM calls are required for a single user interaction.
- Token Throughput: The speed at which a model generates text, measured here in tokens per second (TPS).
The Latency Problem in AI Production
The primary obstacle preventing AI agents from moving beyond "cool demos" into robust production systems is not model intelligence, but latency. Modern customer-facing AI agents typically require three to five LLM calls per interaction. In traditional autoregressive models, these calls are sequential, causing latency to compound and resulting in sluggish, unusable user experiences.
Mercury 2: The Diffusion LLM Approach
Mercury 2 introduces a paradigm shift by utilizing a Diffusion LLM architecture. Unlike standard models that generate one token at a time, Mercury 2 employs parallel refinement.
- Methodology: The model generates a "rough" version of the entire response simultaneously and then refines those tokens in parallel.
- Performance Metrics: This architecture achieves speeds exceeding 1,000 tokens per second (TPS) on standard Nvidia GPUs.
- Comparative Advantage: The speaker claims this is more than five times faster than industry-standard models like Claude 4.5 Haiku and GPT-5 Mini, while operating at a lower cost.
Real-World Applications and Use Cases
The transition from sequential to parallel generation enables use cases that were previously hindered by latency:
- Search and Support: Search Blocks utilizes Mercury 2 to power high-speed search and customer support interactions.
- Real-time Data Processing: WhisperFlow leverages the model for real-time transcript cleanup.
- Voice Avatars: The model meets the critical "sub-second response" requirement necessary for natural-sounding, real-time voice AI avatars.
Integration and Compatibility
A significant barrier to adopting new AI models is the technical overhead of integration. Mercury 2 addresses this by being OpenAI API compatible. This allows developers to swap out existing models in their current tech stack with minimal friction, enabling an immediate performance upgrade without requiring a complete rewrite of the application architecture.
Conclusion
The shift from autoregressive generation to parallel diffusion represents a critical evolution for AI agents. By solving the compounding latency issue, Mercury 2 provides the necessary speed to transform AI from a experimental tool into a reliable, production-ready system. The combination of high throughput, cost-efficiency, and API compatibility positions it as a viable solution for high-demand, real-time AI applications.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Stanford CS153 Frontier Systems | Building the Frontier Ecosystem
Stanford Online

Deterministic Infra for Non-Deterministic AI Agents - Nishant Gupta, Meta Superintelligence Labs
AI Engineer

'Things are going to be okay, in Canada and the U.S.': Thorne
BNN Bloomberg

'No where near normal' but 30-40 oil tankers passing through the Strait 'is better than 0': Mulberry
BNN Bloomberg

'Alphabet has such a dominant position they will be a leader in this space for many years': Clare
BNN Bloomberg

I'M OUT: The $11 Trillion AI Bubble is Breaking!
Steven Van Metre

South Korea bets big on AI with nearly a trillion dollars of investment • FRANCE 24 English
FRANCE 24 English