Inception Labs says its diffusion LLM is 10x faster than Claude, ChatGPT, Gemini
By The New Stack
Key Concepts
- Diffusion Models: A class of generative models that create data by iteratively refining noise into a coherent output (e.g., images or text).
- Autoregressive Models: The standard architecture for LLMs (like GPT), which generates text sequentially, one token at a time, from left to right.
- Inference Latency: The time taken for a model to generate a response; a critical metric for real-time applications like voice agents and coding assistants.
- Parallel Processing: The ability of diffusion models to generate multiple tokens simultaneously, leveraging the architecture of GPUs.
- Denoising/Refinement: The process in diffusion models where a neural network is trained to identify and fix "mistakes" in a rough approximation of the output until it reaches high quality.
- Compute-Bound vs. Memory-Bound: A shift in computational bottlenecks; autoregressive models are often memory-bound (moving weights), while diffusion models are compute-bound (performing arithmetic).
1. Main Topics and Key Points
The discussion centers on the transition from traditional autoregressive LLMs to Diffusion Language Models, pioneered by Inception Labs.
- The Breakthrough: Unlike autoregressive models that predict the "next token," diffusion models start with a rough, noisy approximation of the entire output and iteratively refine it.
- Performance: The primary advantage is speed. Because diffusion models can process multiple tokens in parallel, they can be 5–10x faster than autoregressive models of equivalent quality.
- Efficiency: By shifting from sequential computation to parallel processing, these models better utilize the arithmetic capabilities of modern GPUs.
2. Real-World Applications
Stefano Airmon highlights specific use cases where latency is the primary bottleneck:
- Voice Agents: Real-time interaction requires sub-second response times.
- Coding Assistants: IDE integrations that refactor code or provide suggestions in real-time.
- Search/Information Retrieval: Systems that need to rewrite queries and rerank results rapidly.
- Agentic Workflows: Complex chains where multiple LLMs call tools and interact; speed is critical to prevent compounding delays.
3. Methodologies and Frameworks
- Training Process: Instead of predicting the next word, the model is trained to "fix mistakes." It takes clean text, introduces noise, and learns to reconstruct the original.
- Inference Process: The model generates a full, rough answer and refines it over a small number of "diffusion steps."
- Backwards Compatibility: Inception Labs designed their API to be OpenAI-compatible, allowing developers to switch models by changing only two lines of code.
4. Key Arguments
- The "Autoregressive Bottleneck": Airmon argues that autoregressive models are inherently limited by sequential computation, which cannot be fully optimized due to physical memory bandwidth limits.
- Diffusion as the Future: Airmon posits that diffusion models will eventually become the industry standard because they offer a superior "tokens per dollar per watt" ratio.
- The Role of Reinforcement Learning (RL): RL is used to optimize the model's ability to converge faster and improve reasoning, leveraging the speed of the diffusion inference engine to generate "rollouts" more efficiently.
5. Notable Quotes
- "A typical autoregressive LLM is generating text left to right, one token at a time... it is a fancy autocomplete at the end of the day." — Stefano Airmon
- "In a diffusion language model, you're kind of like getting the GPU to process many tokens at the same time... that's why you can be much more efficient." — Stefano Airmon
- "Eventually, it's all going to be an inference game. The technology that scales better at inference time is the one that is going to win." — Stefano Airmon
6. Data and Research Findings
- 2024 Research Paper: Airmon’s lab published a paper (which won the Best Paper Award at ICML) demonstrating that diffusion models could match the quality of autoregressive models while being 10x faster.
- Mercury 2: The latest model released by Inception Labs, which matches the quality of "speed-optimized" models from frontier labs (like Claude Haiku) while maintaining significantly lower latency.
7. Synthesis and Conclusion
The interview establishes that while autoregressive models currently dominate the LLM landscape, they face a "diminishing returns" wall regarding speed and efficiency. Inception Labs is positioning diffusion models as the high-performance alternative. By treating text generation as an iterative refinement process rather than a sequential prediction task, they have unlocked significant gains in latency. While not yet matching the "frontier" models in absolute intelligence, the focus on speed, cost-efficiency, and developer-friendly integration makes diffusion models a disruptive force in the next generation of AI infrastructure.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Deterministic Infra for Non-Deterministic AI Agents - Nishant Gupta, Meta Superintelligence Labs
AI Engineer

'No where near normal' but 30-40 oil tankers passing through the Strait 'is better than 0': Mulberry
BNN Bloomberg

'Alphabet has such a dominant position they will be a leader in this space for many years': Clare
BNN Bloomberg

Forget Elon’s Data Centers In Space. This Startup Wants To Float Them At Sea
Forbes

Yahoo Finance Live: Daily Market Coverage - June 29, 2026 9AM-11AM (ET)
Yahoo Finance

Why F1 Teams are Replacing Wind Tunnels with Smart Tape | E2305
This Week in Startups

Everyone's Buying AI. Smart Investors Are Buying This Instead. - Robert Kiyosaki
The Rich Dad Channel