Introduction to LLM serving with SGLang - Philip Kiely and Yineng Zhang, Baseten

AI EngineerAbout 5 min readJul 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

SG Lang, LLM serving, VLM, TensorRT-LLM, CUDA graph, speculative decoding, Eagle 3, token acceptance rate, model deployment, performance optimization, open-source contribution, community involvement, base 10, inference stack, model ID, trust, L4 GPU, H100, H200, Blackwell, LM eval, flash infer, quantization, batch size, model server, draft model, target model, token speculation, code base tour, SGLang runtime, domain-specific front-end language, optimized kernels, community Slack, GitHub issues, good first issue, help wanted, development roadmap, SG kernel, SG router, disagre PD disagregation, constraint coding, function calling, open eye compatible server, model architecture, client-server applications, base 10 inference stack, VLM compatibility.

Introduction to LLM Serving with SG Lang

Introduction

The workshop focuses on getting participants comfortable with SG Lang, a fast serving framework for large language models (LLMs) and large vision models (VLMs). The goal is to help users understand how to deploy, optimize, and contribute to SG Lang.

What is SG Lang?

SG Lang is an open-source, high-performance serving framework for LLMs and VLMs. It is used alongside VLM or TensorRT-LLM. Key advantages include:

  • Performance: Excellent performance across various GPUs.
  • Production Ready: Out-of-the-box production readiness.
  • Day Zero Support: Immediate support for new model releases (e.g., Quen, DeepSeek).
  • Community: Strong open-source community.

Who Uses SG Lang?

SG Lang is used by:

  • Base 10 (as part of their inference stack).
  • XAI (for their Glock models).
  • Inference providers.
  • Cloud providers.
  • Research labs.
  • Universities.
  • Product companies (e.g., Koser).

History of SG Lang

  • Archive paper released in December 2023 (18 months ago).
  • Rapid growth to 15,000 GitHub stars.
  • Supports numerous companies and models.
  • Growing international community.

Setting Up and Deploying Your First Model

The workshop uses base 10 for GPU access (L4 GPUs). The process involves:

  1. Using trust to package SG Lang dependencies and commands into a YAML file.
  2. Deploying the packaged configuration to a GPU.
  3. Understanding SG Lang launch server commands, which are essentially a series of flags and configuration options.
  4. Using the model ID from the base 10 UI to set up the URL for calling the model.

The core concept is that SG Lang is essentially a model server that is configured via command-line flags.

Optimizing Performance

CUDA Graph

  • Concept: CUDA graph is a feature that can improve decoding performance.
  • Flag: CUDA graph max BS (batch size).
  • Default: The default max CUDA graph size on L4 GPUs is 8.
  • Issue: When the number of running requests exceeds the max CUDA graph size, CUDA graph is disabled, reducing performance.
  • Solution: Adjust the CUDA graph max BS parameter to a higher value (e.g., 32) to accommodate realistic batch sizes during benchmarking.
  • Benchmarking: Use LM eval to send requests and analyze server logs to determine if CUDA graph is enabled during decoding.
  • Command Example: The LMEval command line is shared to evaluate the model performance. It specifies the model, URL, port, number of concurrent requests, max generation tokens, and the evaluation dataset (GSMK).

Eagle 3 Speculative Decoding

  • Concept: Eagle 3 is a speculative decoding algorithm that can improve performance by speculating on future tokens.
  • Key Parameters:
    • speculative decoding algorithm: Set to eagle.
    • draft model path: Path to the draft model (derived from the target model).
    • numbers depths: Depth of the drafting.
    • eagle top K: Top K value for token selection.
    • draft verify tokens: Number of tokens to verify.
  • Tuning: A script is provided to tune the numbers depths, eagle top K, and draft verify tokens parameters.
  • Benchmarking: It's crucial to use prompts representative of the actual workload to ensure accurate speculation and parameter tuning.
  • Draft Model: Eagle 3 differs from standard draft-target approaches by using multiple layers of the target model to build the draft model.

Community and Getting Involved

  • GitHub: Star the project, file issues, and bug reports.
  • Tagging System: Use the tagging system to find issues to contribute to (e.g., "good first issue," "help wanted").
  • Twitter: Follow SG Langis.org on Twitter.
  • Slack: Join the community Slack for online and in-person meetups (slack.sglang.ai).

Codebase Overview

The SG Lang codebase consists of:

  • SGLang Runtime: The core runtime engine.
  • Domain-Specific Front-End Language: A language for defining and controlling model behavior.
  • Optimized Kernels: High-performance CUDA kernels for attention, normalization, activation, and GEMM operations.

Key components include:

  • SG Kernel: Implements attention, normalization, activation, and GEMM operations.
  • SG Router: Handles cashware rooting.
  • SRT (Python): The core Python component that supports disagre PD disagregation, constraint coding, function calling, and an open eye compatible server.

Contributing to the Codebase

  1. Use SG Lang to identify missing features or issues.
  2. Raise a new issue on GitHub.
  3. Look for issues tagged as "good first issue" or "help wanted."
  4. Refer to the development roadmap for planned features.
  5. Contribute to specific components based on expertise (e.g., CUDA kernels, routing, Python runtime).

Conclusion

SG Lang is a powerful and rapidly evolving framework for LLM serving. Its performance, flexibility, and strong community make it a compelling choice for both research and production deployments. By understanding the core concepts, optimization techniques, and contribution pathways, users can effectively leverage SG Lang to build and deploy high-performance LLM applications.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.