NVIDIA Monopoly is DEAD | OPEN-SOURCE Chips Are HERE!
By Hefty LLM
Overview of Tenstorrent’s Architectural Shift
Jim Keller, the legendary chip architect behind AMD’s Zen, Apple’s A-series, and Tesla’s FSD silicon, has spent the last four years developing a new AI chip architecture at Tenstorrent. This architecture aims to disrupt Nvidia’s dominance by rejecting traditional GPU design principles, focusing on efficiency, and utilizing open-source software to reduce inference costs by a factor of five.
Core Architectural Innovations
Tenstorrent’s approach is built on the premise that AI workloads are highly predictable, unlike the random nature of video game rendering.
- Removal of Hardware Overhead: Traditional GPUs dedicate significant silicon area to "hidden overheads" like hardware schedulers, cache controllers, and memory management units. Tenstorrent removed these entirely.
- Compiler-Driven Execution: Because AI math is predictable, the compiler maps the entire journey of data before the chip begins computing. This shifts the intelligence from the hardware to the software.
- Ten6 Cores: The architecture uses tiles containing five independent RISC-V cores, each with its own local SRAM. This decentralized design ensures that no core remains idle waiting for global synchronization, as there is no global clock forcing cores to work in lockstep.
- Memory Strategy: Instead of expensive, high-bandwidth HBM (High Bandwidth Memory), Tenstorrent uses standard GDDR6. By using the compiler to "pre-fetch" data into the 200MB of on-chip SRAM just before it is needed, they bypass the bandwidth limitations of cheaper memory.
Scaling and Networking
To solve the "scaling problem" where chips spend more time communicating than computing, Tenstorrent integrated networking directly into the silicon:
- Integrated Ethernet: Every "Blackhole" chip features 400 Gbps Ethernet baked directly into the hardware, allowing the chip to act as both a processor and a router.
- Unified Fabric: The compiler pre-maps data movement across the entire cluster. This allows 32 chips in a "Galaxy" server to function as a single, unified brain, scaling to over a thousand chips without the latency or heat issues associated with external InfiniBand switches.
Performance and Economic Impact
- Efficiency: On models like DeepSeek-R1, the architecture achieves 350 tokens per second.
- Cost: The cost of operation is approximately $6 per million tokens, compared to $30 per million tokens on Nvidia hardware.
- Open Source: Unlike Nvidia’s proprietary CUDA ecosystem, Tenstorrent’s software stack is fully open-source, allowing the global developer community to contribute to its optimization.
Challenges to Adoption
Despite the technical advantages, two primary barriers prevent immediate mass-market adoption:
- The "10% Problem": While 90% of Hugging Face models run out of the box, enterprise clients (hospitals, financial institutions) require 100% reliability. The software stack for the new "Blackhole" chips is still maturing compared to the established, battle-tested software of previous generations.
- The "Jim Keller Pattern": Historically, Keller has left companies (AMD, Apple, Tesla) shortly after building the foundation. However, this is the first time he is serving as CEO rather than a hired architect, suggesting a long-term commitment to the company’s success.
Key Concepts
- Inference: The process of running a trained AI model to make predictions or generate content.
- RISC-V: An open-standard instruction set architecture (ISA) that allows for custom, efficient processor design.
- SRAM (Static Random Access Memory): High-speed, low-latency memory located directly on the chip.
- HBM (High Bandwidth Memory): A high-performance, expensive memory interface used in most modern AI GPUs.
- GDDR6: A type of synchronous graphics RAM designed for high-bandwidth applications, typically found in gaming hardware.
- Compiler: The software that translates high-level code into machine-specific instructions; in Tenstorrent’s case, it also manages data orchestration.
- NVLink: Nvidia’s proprietary high-speed interconnect technology for linking multiple GPUs.
- InfiniBand: A high-speed network communications standard used in high-performance computing.
- Blackhole: The current generation of Tenstorrent chips featuring integrated Ethernet.
- CUDA: Nvidia’s proprietary parallel computing platform and programming model.
Synthesis
Tenstorrent represents a fundamental shift from "brute force" hardware design to "software-defined" silicon. By leveraging open-source architecture and eliminating the overhead of traditional GPU management, Jim Keller has created a system that is significantly cheaper and more efficient for AI inference. While software maturity remains a hurdle, the shift toward an open-source, community-driven ecosystem provides a faster trajectory for improvement than traditional proprietary models.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

99% Follow Goals, Only 1% Do this
Him-eesh Madaan

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

¿Trabajas en Oficina? EL ERROR que comete el 99% con Julieta Manzano | Martha Debayle
Martha Debayle

How East India Company Captured India | Nitish Rajput | Hindi
Nitish Rajput @

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED