Key Concepts:
- Inference Compute: Applying knowledge acquired by an AI model to solve problems or complete tasks, analogous to humans applying their education in the workforce.
- Pre-training: The initial general acquisition of skills and knowledge for an AI model.
- Agentic AI: AI systems comprised of multiple agents that collaborate to perform complex tasks, requiring significant compute power.
- Ultra-Fast Inference: Inference with minimal latency, crucial for interactive applications and machine-to-machine communication in agentic workflows.
- Heterogeneous Computing: Utilizing different types of specialized hardware to optimize for various inference workload profiles (e.g., cost, latency, throughput).
- Chiplets: Small, modular chips that can be co-packaged to create larger, more flexible computing platforms.
- 3D Stacking: A packaging technique that involves vertically stacking memory directly on top of compute units to minimize the distance data must travel.
- Reasoning/Inference Time Compute/Test Time Compute: Models that are beginning to think more, which feeds directly into the agentic era.
1. Introduction to Dmatrix and Inference Compute
- Dmatrix is an inference compute company focused on building solutions for generative AI workloads, emphasizing ultra-fast and commercially viable inference.
- Founded in 2019, Dmatrix concentrates solely on inference compute, not pre-training, HPC, or graphics workloads.
- The company's thesis is that inference computing will become the largest computing workload.
2. Defining Inference in AI
- Inference is when a model takes data and uses it to solve a problem or complete a task.
- Analogy to human learning: Pre-training is like education, and inference is like applying knowledge in the workforce.
- Fine-tuning models for specific expertise is analogous to humans becoming domain experts.
- RAG (Retrieval Augmented Generation) is like accessing external information when a model lacks specific knowledge, similar to humans using the internet.
3. Challenges of Inference
- Challenges are primarily software-focused: integrating models into software workflows and developing user interfaces for enterprise users.
- Broad-based adoption is hindered by the need for workforce retraining and significant human intervention (prompt engineering).
- Agentic AI promises to simplify task execution by enabling users to describe tasks at a high level, with agents breaking them down and collaborating.
4. Dmatrix's Solution for Agentic Workflows
- Dmatrix aims to address the escalating compute costs and latency concerns associated with agentic AI.
- The company's platform is designed to support multiple users and agents without sacrificing per-user latency, enabling optimized agentic experiences.
5. Infrastructure Needs for Inference vs. Training
- Inference requires different infrastructure than pre-training, as it is not a "one-size-fits-all" scenario.
- User profiles vary significantly, with some prioritizing cost, others interactivity (latency), and others throughput.
- Dmatrix believes the world of inference will be heterogeneous, with dedicated hardware for specific needs.
6. Distance and Latency
- The distance data has to travel to compute is a key challenge.
- Minimizing this distance is crucial for generative AI workloads due to frequent memory access for caching.
- Dmatrix's platform collocates compute and memory to minimize data travel distance.
7. Dmatrix's Chiplet-Based Architecture and 3D Stacking
- Dmatrix uses chiplets (smaller chips) co-packaged into a fabric, providing elasticity and modularity.
- Instead of keeping compute and memory close, Dmatrix overlays memory directly on top of compute using 3D stacking.
- This allows data to "rain down" into the compute, increasing surface area and minimizing distance.
8. Target Customers and Use Cases
- Dmatrix targets anyone using generative AI who wants fast inference, including enterprises, clouds, neoclouds, and sovereigns.
- Interactive inference is becoming increasingly important as users prefer applications with quick response times.
9. AI Infrastructure Maintenance and Skills
- Organizations need to retool and reskill their workforce for AI infrastructure.
- The existing general-purpose computing infrastructure is being augmented or replaced with accelerated computing.
- There is a need for orchestration software to route queries to appropriate resources in heterogeneous, potentially distributed, fleets.
10. Dmatrix Roadmap
- Focus on releasing the 3D capability product.
- First-generation product: PCI card with eight chiplets packaged into an accelerator unit with a DMX bridge, already shipping to customers.
- Next product will incorporate 3D stacking.
- Dmatrix is focusing on reasoning/inference time compute as the future direction.
11. Product Details
- Current Product: A PCI card rated at 600W with two chips, each containing four chiplets (total of eight chiplets).
- Accelerator Unit: Contains two of the PCI cards bridged with a DMX bridge for all-to-all chiplet communication. Treated as a single unit by software.
12. Upcoming 3D Product
- The upcoming product will feature 3D stacking.
- Demo of 3D product capability in the next few months with product to market about a year later.
Synthesis/Conclusion:
Dmatrix is positioned as a key player in the evolving landscape of AI inference, particularly with the rise of agentic AI. Their focus on ultra-fast, cost-effective, and highly interactive inference solutions, driven by innovative chiplet architecture and 3D stacking, addresses critical challenges in deploying AI at scale. By minimizing data travel distance and optimizing compute utilization, Dmatrix aims to empower a wide range of customers, from enterprises to sovereigns, seeking to leverage generative AI effectively. The company is actively developing its 3D product and preparing for the future of reasoning-based AI models, emphasizing their commitment to enabling efficient and scalable inference compute.
AI summaries can miss context or contain errors. Check important details against the original video.





