Stanford CS25: Transformers United V6 I Serving Transformers: Lessons from the Trenches

Stanford OnlineAbout 3 min readJun 5, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • Inference Engineering: The process of deploying and optimizing AI models for production, focusing on latency, throughput, and cost-efficiency.
  • Workload Archetypes: Chatbot+, Background Agents, and Data Processors.
  • SLAs/SLOs: Service Level Agreements/Objectives defining latency (Time to First Token, Time per Output Token) and throughput requirements.
  • Inference Engines: Software frameworks (e.g., vLLM, SGLang, TensorRT-LLM) that orchestrate GPU execution.
  • Speculative Decoding: A technique using a smaller "draft" model to predict tokens, which are then verified by a larger target model to increase throughput.
  • Quantization: Reducing model precision (e.g., FP8, FP4) to lower memory bandwidth requirements and increase speed.
  • Observability: The ability to debug system failures and performance bottlenecks using logs and metrics.

1. Application Archetypes and Workload Definition

Charles categorizes AI applications into three distinct archetypes, each with unique engineering constraints:

  • Chatbot+: Human-interactive, requires low latency (Time to First Token).
  • Background Agents: Asynchronous, human-in-the-loop but not waiting for immediate output; latency tolerance is higher (seconds to hours).
  • Data Processors: High-volume, bursty, high latency tolerance; focus is on "mega-tokens per dollar."

Workload Metrics:

  • QPS (Queries Per Second): User-driven, often seasonal/variable.
  • Token Counts: Input (pre-fill) and output (decode) lengths.
  • Prefix Reuse: Caching previous computations to save GPU resources.
  • Latency Budgets: Defined by the user experience (Time to First Token vs. Inter-token latency).

2. Inference Engines and Hardware

The "stack" involves coordinating CPU-based scheduling with GPU-based execution.

  • Engines:
    • vLLM: Wide adoption, open governance, enterprise-ready.
    • SGLang: Performance-focused, "startup-y" culture, aggressive optimization.
    • TensorRT-LLM: Compiled C++ runtime, best for small models/batch sizes.
  • Hardware:
    • Data Center GPUs (SXM form factor): Essential for high-bandwidth memory (HBM) and power delivery.
    • Tensor Cores: Specialized units for matrix multiplication; critical for performance.
    • Bottlenecks: Decode phases are memory-bandwidth bound, while pre-fill phases are compute-bound.

3. Performance Optimization Framework

Charles suggests a tiered approach to optimization:

  1. Big Wins: Speculative decoding (using draft models) and quantization (FP8/FP4).
  2. Host Engineering: Reducing CPU overhead (e.g., CUDA graph capture, caching pointers in Python).
  3. GPU Kernel Optimization: Low-level tuning (e.g., Nsight Compute) for marginal gains at scale.

Key Quote: "Speculation is all you need... it gives you these very large integral measured speed-up factors."

4. Observability and Debugging

  • Logging: Log token IDs, not just strings, to identify tokenizer bugs.
  • Metrics: Track P50, P95, and P99 latencies. Tail latencies are often caused by queuing delays rather than model compute time.
  • Tools: Use pi-spy for host-side bottlenecks and Nsight Systems for GPU-side profiling.
  • Evals: Treat evaluations as "tests." They are model-agnostic and essential for comparing performance across different model versions or quantization levels.

5. Deployment and Scaling

  • Multi-tenancy: Avoid spinning up machines from scratch. Maintain a buffer of idle replicas to handle traffic spikes.
  • Lazy Loading: Load file systems and dependencies concurrently with container startup to reduce cold-start times.
  • Failure Handling: Assume hardware failure (e.g., H100s fail in days/weeks). Build systems with redundancy so that individual replica failures do not crash the entire service.

Synthesis and Conclusion

Inference is the "revenue center" of AI, whereas training is the "cost center." As models become more commoditized, the competitive advantage shifts to inference efficiency. The future of the field lies in "lossy" optimizations (pruning, layer skipping), custom speculative draft models, and the integration of AI agents into the software engineering lifecycle itself. Engineers should prioritize empirical measurement (benchmarking) and observability over premature low-level kernel optimization.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.