How to Build Your Own AI Data Center in 2025 — Paul Gilbert, Arista Networks

AI EngineerAbout 5 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Backend Network (GPU interconnect)
  • Frontend Network (Storage access)
  • Inference Networks
  • RDMA (Remote Direct Memory Access)
  • RoCEv2 (RDMA over Converged Ethernet version 2)
  • PFC (Priority Flow Control)
  • ECN (Explicit Congestion Notification)
  • Collectives (Nickel)
  • Lossless Ethernet
  • Cluster Load Balancing
  • Ultra Ethernet Consortium
  • Job Completion Time
  • Scale-up vs. Scale-out
  • Wire Rate
  • Entropy Load Balancing
  • Smart System Upgrade
  • Telemetry and Visibility

AI Network Infrastructure: A Deep Dive

Introduction

Paul Gil, a tech lead at Arista Networks, discusses the infrastructure required for training and inferencing AI models, focusing on the networking aspects. He emphasizes the differences between traditional enterprise networks and the specialized networks needed for AI workloads, particularly concerning bandwidth, latency, and traffic patterns.

Training vs. Inference

  • GPU Requirements: Training typically requires significantly more GPUs than inference. Dr. Wed Sosa's comparison shows training potentially needing 18 times the GPU power of inference, although this is changing with Chain of Thought and reasoning models.
  • Model Size and Infrastructure: A model trained on 248 GPUs for 1-2 months might only require four H100s for inference after fine-tuning and alignment.
  • LLM Inference Growth: LLMs are now a significant portion of inference workloads, requiring more robust infrastructure.

Network Architecture

  • Backend Network (GPU Interconnect): This network connects GPUs and is isolated due to the high cost and power consumption of GPUs. It typically consists of eight GPUs per server connected to high-speed leaf and spine switches. The backend network operates at 400 Gbps or higher.
  • Frontend Network (Storage Access): This network provides access to storage for training data. While important, it's generally less demanding than the backend network.
  • Isolation: AI networks are often completely isolated from other enterprise networks to ensure performance and security.

Hardware and Software Considerations

  • GPUs: Modern GPUs like the H100 have eight 400 Gbps ports for GPU interconnect and additional Ethernet ports for frontend network connectivity.
  • Scale-Up vs. Scale-Out: While GPUs themselves are generally not scalable up (cannot add more GPUs to a server), the network is designed to scale out by adding more GPU servers.
  • Software: Cuda and Nickel are key software components. Nickel's collective communication patterns influence network traffic.
  • Data Center Applications vs. AI Networks: Traditional data center applications are more fault-tolerant with load balancing and failover mechanisms. AI networks are more sensitive; a single GPU failure can impact the entire training job.

Traffic Patterns and Bandwidth

  • Bursty Traffic: AI workloads generate bursty traffic as all GPUs may transmit simultaneously at high bandwidth (up to 400 Gbps per GPU).
  • No Oversubscription: AI networks require a 1:1 subscription ratio to avoid congestion and packet loss.
  • High Bandwidth Requirements: A single H100 server can potentially generate 9.6 terabits per second of traffic (8 x 400 Gbps + 4 x 400 Gbps).
  • East-West vs. North-South Traffic: AI networks have significant east-west traffic (GPU-to-GPU) in addition to north-south traffic (storage access). East-west traffic is particularly demanding.

Load Balancing and Congestion Control

  • Entropy Load Balancing Limitations: Traditional entropy load balancing (based on five tuples) is insufficient for AI networks because GPU traffic often originates from a single IP address.
  • Bandwidth-Aware Load Balancing: Arista uses tools that load balance based on the percentage of bandwidth utilization on uplinks, achieving up to 93% utilization.
  • RoCEv2: RoCEv2 is used for congestion control, employing PFC (Priority Flow Control) for emergency stops and ECN (Explicit Congestion Notification) for gradual slowdowns.

Infrastructure Challenges

  • Optics and Cabling: Building networks with thousands of GPUs introduces challenges related to optics, transceivers, and cable reliability.
  • Power: AI servers consume significant power (e.g., 10.2 kW for an H100 server with eight GPUs), requiring high-density racks (100-200 kW) and water cooling.
  • Buffering: Switches need adequate buffering to handle bursty traffic.

Network Design Principles

  • Simplicity: Keep the network design as simple as possible to minimize potential points of failure.
  • Isolation: Isolate the AI network from other networks for performance and security.
  • Address Space: Use point-to-point addressing (e.g., /30 or /31 subnets) and consider IPv6 if IPv4 address space is limited.
  • BGP: Use BGP as the routing protocol for its simplicity and speed.
  • EVPN/VXLAN: Use EVPN/VXLAN for multi-tenancy environments.

Arista Solutions

  • EOS (Arista Operating System): Arista's EOS offers features for building AI networks, including lossless Ethernet, adjustable buffers, and monitoring tools.
  • RDMA Packet Capture: Arista switches can capture RDMA packets and headers to diagnose packet loss issues.
  • AI Agent: An AI agent running on GPUs communicates with the switch to verify configuration and provide statistics on packet flow and RDMA errors.
  • Smart System Upgrade: Arista's smart system upgrade allows upgrading switch software without interrupting GPU workloads.
  • Cluster Load Balancing: Load balancing based on the collective being run.

Future Trends

  • 800 Gbps and Beyond: Networks are moving to 800 Gbps, with 1.6 Tbps expected in the near future.
  • Ultra Ethernet Consortium: The Ultra Ethernet Consortium is developing new Ethernet standards to improve congestion control and network performance for AI workloads. Version 1.0 is expected in Q1 2025.

Conclusion

Building AI networks requires a different approach than traditional enterprise networks. Key considerations include high bandwidth, low latency, lossless transport, robust congestion control, and specialized monitoring tools. Arista Networks offers solutions and best practices to address these challenges and enable the deployment of large-scale AI infrastructure. The focus is on minimizing job completion time and providing visibility into network performance to ensure the efficient training and inferencing of AI models.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.