Model types and performance bottlenecks
By Google Cloud Tech
Key Concepts
- AI Model Architectures: Large Language Models (LLMs), Diffusion Models, Visual Language Models (VLMs), Mixture of Experts (MoE).
- Performance Bottlenecks: Compute-bound, Memory-bound, Memory Bandwidth-bound, Network-bound.
- Optimization Techniques: Quantization, Tensor Parallelism, Pipeline Parallelism, Model Compilation (e.g., Tensor RT).
- Hardware Components: TPUs (Tensor Processing Units), VRAM (Video Random Access Memory), Accelerators, Interconnects.
Introduction to AI Model Diversity and Bottlenecks
The performance of AI models during inference is not uniform and depends heavily on their architecture. Understanding the specific type of model is crucial for identifying and resolving performance limitations. This video explores four common bottlenecks: compute, memory, memory bandwidth, and networking.
Four Common AI Model Architectures
-
Large Language Models (LLMs):
- Description: Trained on vast text datasets to understand and generate language.
- Defining Feature: Billions, sometimes trillions, of parameters.
- Examples: Gemma, Lama, Deepseek, FE.
-
Diffusion Models:
- Description: Used in text-to-image generators. They start with noise and progressively denoise it to match a prompt and create an image.
- Examples: Dolly, Imagine, Stable Diffusion.
-
Visual Language Models (VLMs):
- Description: Multimodal models that can process and understand both images and text. They can describe images or answer questions about them.
- Examples: Gemini (Google), ChatGPT.
-
Mixture of Experts (MoE):
- Description: Composed of many smaller "expert" models. For any given input, only a subset of these experts are activated.
- Examples: Kimmy K2, Mixrol.
Four Common Performance Bottlenecks
-
Compute-bound:
- Definition: The processor (e.g., TPU pod) is running at 100% utilization, performing calculations as fast as possible.
- Common Scenarios: Training dense models, inference steps of diffusion models where calculations are intensive.
- Solution: Faster chips or more chips. Adding more chips is often easier than other optimization methods.
-
Memory-bound:
- Definition: The model's parameters, activations, and data batch exceed the available VRAM on the accelerator. The issue is insufficient space, not speed.
- Common Scenarios: Giant LLMs. A 175 billion parameter model can require over 350 GB of VRAM for storage alone. If a TPU has only 80 GB, it cannot even begin processing.
- Solution: Provisioning machines with accelerators that have higher VRAM capacity.
-
Memory Bandwidth-bound:
- Definition: The processor is underutilized (e.g., 50% utilization) because it is waiting for data to be transferred from VRAM to the processing cores. The bottleneck is the speed of data transfer.
- Common Scenarios: LLM inference where parameters are constantly moved to the processor, and MoE models where experts are frequently swapped in and out of memory.
- Solution: Optimizing data movement between VRAM and cores.
-
Network-bound:
- Definition: Occurs when scaling to multiple servers. The model is split across machines, and progress is halted until all servers complete their tasks and communicate results.
- Common Scenarios: Distributed training of LLMs, large-scale models requiring communication between experts on different machines.
- Solution: High-speed, low-latency interconnects between servers.
Practical Manifestations and Solutions by Model Type
-
Large Language Models (LLMs):
- Primary Bottleneck: Memory capacity.
- Platform Engineer Solution: Provision machines with multiple high VRAM accelerators.
- AI Engineer Solution: Quantization. Reducing model weights from 16-bit floats to 8-bit or 4-bit integers can significantly reduce memory usage.
- Training Bottleneck: Network wall. Distributed training (tensor or pipeline parallelism) requires massive synchronization. Slow interconnects lead to idle TPUs.
-
Diffusion Models:
- Primary Bottleneck: Pure compute.
- Reason: The denoising loop involves numerous forward passes through computationally intensive units (e.g., convolutions) without significant data transfer.
- Solution: Raw power (faster accelerators) or model compilation using frameworks like Tensor RT to optimize the computational graph.
-
Mixture of Experts (MoE):
- Bottlenecks: Trade compute for memory bandwidth and networking.
- Mechanism: The gating network selects experts, requiring their weights to be loaded into cores.
- If Experts on Same Chip: Limited by High Bandwidth Memory (HBM) speed.
- If Experts on Different Machines: Severely network-bound, waiting for remote expert communication.
- Solution: Ultra-fast, low-latency interconnects are critical for success with very large MoE models.
Conclusion and Key Takeaways
Understanding a model's architecture is not just about predicting potential issues but also about developing a strategic approach to hardware selection, parallelization strategies, and software optimizations. The four primary bottlenecks—compute, memory capacity, memory bandwidth, and inter-server network—affect different model types in distinct ways. When encountering performance issues, it's essential to diagnose whether the system is waiting for computation, data, or experiencing saturation in memory bandwidth or network communication, rather than solely blaming the hardware.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.

