Hacking the Inference Pareto Frontier - Kyle Kranen, NVIDIA

AI EngineerAbout 6 min readAug 3, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Pareto Frontier: A visualization of the trade-offs between different performance metrics (quality, latency, cost).
  • Disaggregation: Separating the prefill and decode phases of LLM inference onto different workers.
  • KV Caching: Caching key and value vectors to avoid recomputing them during auto-regressive generation.
  • Inference Time Scaling: Re-querying the model multiple times to improve quality.
  • Worker Specialization: Using different types of workers (aggregated vs. disaggregated) based on workload characteristics.
  • Dynamic Load Balancing: Adjusting the number of prefill and decode workers in real-time to adapt to changing user demand.

1. Pareto Frontier and Application-Specific Requirements

  • The speaker, Kyle Cranon, emphasizes that a good model and system, tailored to deployment constraints, are crucial for success.
  • Three key factors for deployability: quality (accuracy), latency (speed), and cost (per-request expense).
  • The Pareto frontier represents the optimal trade-offs between these factors. The speaker notes that the frontier is hard to plot in 3D, so he will show it in 2D.
  • The ideal operating point on the Pareto frontier depends on the application.
    • Example: Personal cancer cures prioritize quality over latency and cost.
    • Example: Tab completion in IDEs demands low latency.
    • Example: Async code commits prioritize quality and cost over latency.

2. Techniques for Manipulating the Pareto Frontier

  • Common techniques and their effects:
    • Quantization: Speeds up latency and reduces cost by enabling higher batch sizes.
    • Retrieve Augmented Generation (RAG): Increases quality but slows down latency and increases cost.
    • Reasoning: Increases quality, latency, and cost due to more token generation.
    • Model Configuration Changes: Can affect speed, cost, and quality (e.g., non-haloed context parallelism).
  • These techniques can be combined to achieve desired trade-offs.
    • Example: RAG to improve quality, followed by quantization to mitigate latency increase.

3. Three Drivers for Pareto Frontier Improvement: Scale, Structure, and Dynamism

  • Scale:
    • Disaggregation: Separating prefill (compute-bound) and decode (memory-bound) phases onto different workers.
      • Benefits: Granular load matching, simplified scheduling (reduced conflicts from inflight batching).
      • Performance: Llama 7B example shows up to 2x tokens per second per GPU at a fixed latency with disaggregation on 16 H100s.
      • Constraints: Low input length use cases see little speedup. Configuration (balance between prefill and decode workers) is critical.
    • Routing: Optimizing the transfer of KV cache between machines.
      • Naive Routing: Random routing or KV-based routing (can lead to queuing).
      • Smart Routing: Minimizes a cost function that includes both prefix match (KV cache hit rate) and node load.
      • As deployment scales, KV cache hit rate increases asymptotically.
  • Structure:
    • Agents: Workloads with moderately predictable usage patterns.
    • Inference Time Scaling: Re-querying the model to improve quality.
      • Example: An 8B model requering multiple times can achieve quality comparable to a 49B model.
      • Trade-off: Increases quality at the cost of speed and cost, but can be more efficient than using a larger model directly.
    • Leveraging structure for better scheduling:
      • Removing round trips by making requeries come from the router.
      • Making the LM scheduler aware of repeat work.
    • KV Manipulation: Offloading KV cache to host memory during tool calls to avoid recomputation.
  • Dynamism:
    • Worker Specialization: Using different worker types (aggregated vs. disaggregated) based on input/output sequence length (ISL/OSL) histograms.
      • Example: Aggregated workers for lower ISL and higher OSL, disaggregated for middle ranges, and disaggregated with context parallelism for long context.
    • Dynamic Load Balancing: Autoscaling prefill and decode workers in real-time to adapt to changes in user usage distribution.
      • Essential for maximizing the potential of disaggregation.

4. Disaggregation in Detail

  • KV caching leads to two phases: prefill (filling the KV cache) and decode (generating new tokens).
  • Disaggregation allows these phases to run on different workers, optimizing resource utilization.
  • Prefill is compute-bound, while decode can be memory-bound.
  • Granular load matching: Using fewer GPUs for prefill with lower batch sizes, and more GPUs for decode with larger batch sizes.
  • Scheduling conflicts: Inflight batching on the same machine can cause scheduling conflicts. Disaggregation simplifies scheduling.
  • Configuration: Balancing the number of prefill and decode workers is crucial.

5. Routing Strategies

  • Naive routing: Randomly routing requests to workers.
  • KV-based routing: Routing requests to workers with matching KV cache, but can lead to overloading.
  • Smart routing: Optimizes for both KV cache hit rate and worker load.
  • As the deployment scales, the KV cache hit rate increases, reducing the amount of prefill work.

6. Inference Time Scaling and KV Manipulation

  • Inference time scaling: Re-querying the model multiple times to improve quality.
  • KV manipulation: Offloading KV cache to host memory during tool calls to avoid recomputation.
  • Example: Offloading KV cache during a 30-second tool call and moving it back to GPU memory when the tool completes.

7. Dynamism and Worker Specialization

  • Worker specialization: Using different worker types based on ISL/OSL histograms.
  • Dynamic load balancing: Autoscaling prefill and decode workers in real-time to adapt to changes in user usage distribution.
  • Example: A change in user distribution can create more demand for prefill workers than decode workers, requiring autoscaling.

8. Conclusion

  • Breaking the Pareto frontier requires a combination of techniques, including disaggregation, routing, inference time scaling, KV manipulation, worker specialization, and dynamic load balancing.
  • The optimal approach depends on the specific application and its requirements.
  • Dynamo, an open-source project from NVIDIA, aims to enable data center scale inference and manipulate the Pareto frontier.

9. Notable Quotes

  • "A good model and a good system that takes into account the actual constraints for what you need from your deployment is actually key to the success of both your deployment and the application that is backed by it." - Kyle Cranon
  • "These techniques can be compounded." - Kyle Cranon

10. Technical Terms

  • KV Cache: Key-value cache, used to store intermediate results during LLM inference.
  • Prefill: The initial phase of LLM inference, where the KV cache is filled.
  • Decode: The phase of LLM inference where new tokens are generated.
  • ISL: Input Sequence Length.
  • OSL: Output Sequence Length.
  • Tensor Parallelism: A technique for distributing a model across multiple GPUs.
  • Context Parallelism: A technique for distributing the context of a model across multiple GPUs.
  • HBM: High Bandwidth Memory, a type of memory used in GPUs.

11. Logical Connections

  • The talk begins by establishing the importance of the Pareto frontier and application-specific requirements.
  • It then introduces various techniques for manipulating the Pareto frontier, emphasizing that these techniques can be combined.
  • The talk then focuses on three key drivers for Pareto frontier improvement: scale, structure, and dynamism.
  • Each of these drivers is discussed in detail, with specific examples and case studies.
  • The talk concludes by summarizing the main takeaways and highlighting the importance of Dynamo, an open-source project from NVIDIA.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.