vLLM: Easily Deploying & Serving LLMs

NeuralNineAbout 5 min readSep 6, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • VLM (vLLM): An easy-to-use and fast Python library for LLM inference and serving.
  • Hugging Face Models: Pre-trained language models available on the Hugging Face platform.
  • Model Serving: Deploying a model to make it accessible via an API.
  • Fine-tuned Models: Models that have been further trained on a specific dataset for a particular task.
  • GGUF: A file format for storing quantized models, often used with llama.cpp.
  • Tokenizer: A component that converts text into numerical tokens that the model can understand.
  • Sampling Parameters: Parameters that control how the model generates text, such as temperature and max tokens.
  • OpenAI API: An API provided by OpenAI for accessing their language models.
  • API Key: A secret key used to authenticate requests to an API.
  • GPU Memory Utilization: The percentage of GPU memory that the model is allowed to use.

1. Using Hugging Face Models with VLM in Code:

  • VLM simplifies the process of using Hugging Face models in Python code.
  • Instead of using the transformers package and manually setting up tokenizers, VLM allows you to directly load and use models with a single line of code.
  • Example:
    from vllm import llm, SamplingParams
    
    lm = llm("TinyLlama/TinyLlama-1.1B-Chat-v1.0")
    sampling_params = SamplingParams(max_tokens=128, temperature=0.7)
    outputs = lm.generate(["What is special about Python programming language?"], sampling_params)
    print(outputs[0].outputs[0].text.strip())
    
  • GPU_memory_utilization parameter can be used to limit the amount of GPU memory used by the model to prevent out-of-memory errors.
    • Example: lm = llm("TinyLlama/TinyLlama-1.1B-Chat-v1.0", gpu_memory_utilization=0.7)

2. Serving Hugging Face Models Locally with VLM:

  • VLM can be used to serve Hugging Face models locally, providing an OpenAI-compatible API.
  • This allows you to interact with the model using the OpenAI Python package.
  • Command to serve a model:
    vllm serve --model TinyLlama/TinyLlama-1.1B-Chat-v1.0 --gpu_memory_utilization 0.7 --api_key neural9key
    
  • The model is hosted at http://localhost:8000/v1.
  • Python code to interact with the hosted model:
    from openai import OpenAI
    
    client = OpenAI(api_key="neural9key", base_url="http://localhost:8000/v1")
    response = client.chat.completions.create(
        model="TinyLlama/TinyLlama-1.1B-Chat-v1.0",
        messages=[
            {"role": "system", "content": "You are a helpful assistant."},
            {"role": "user", "content": "Summarize the idea of docker in simple words"}
        ]
    )
    print(response.choices[0].message.content)
    

3. Serving Custom Fine-tuned Models with VLM:

  • VLM can also be used to serve custom fine-tuned models, including those trained with Unsloth.
  • This requires providing the path to the GGUF file, the tokenizer directory, and the chat template.
  • Command to serve a fine-tuned model:
    vllm serve --model gguf_model_scratch_smaller/unsloth/q4k_m.gguf --tokenizer gguf_model_scratch_smaller --served_model_name tuned_tutorial_model --chat_template gguf_model_scratch_smaller/chat_template.jinger --gpu_memory_utilization 0.7 --api_key neural9key
    
  • The --served_model_name parameter specifies the name that will be used to identify the model in the OpenAI API.
  • Example: The video uses a fine-tuned model that is trained to extract information about a person from a given text.
  • The Python code to interact with the hosted fine-tuned model is similar to the previous example, but the model parameter in the client.chat.completions.create function should be set to the value provided in --served_model_name (e.g., "tuned_tutorial_model").

4. Key Arguments and Perspectives:

  • VLM simplifies LLM deployment and inference, making it accessible to a wider range of users.
  • Serving models locally with an OpenAI-compatible API allows for easy integration with existing tools and workflows.
  • VLM enables users to deploy their own custom fine-tuned models, giving them more control over the model's behavior and capabilities.

5. Notable Quotes:

  • "VLM is an easy-to-use and fast Python library that allows us to do LLM inference and serving."
  • "You can do that with any model provided that you have the necessary hardware to run this on your system."
  • "This will run the model and host it like an OpenAI model."

6. Technical Terms:

  • LLM (Large Language Model): A type of artificial intelligence model that is trained on a large amount of text data and can generate human-like text.
  • Inference: The process of using a trained model to make predictions on new data.
  • Serving: The process of deploying a model to make it accessible via an API.
  • Quantization: A technique for reducing the size of a model by reducing the precision of its weights.
  • GGUF: A file format for storing quantized models, often used with llama.cpp.
  • Tokenizer: A component that converts text into numerical tokens that the model can understand.
  • Sampling Parameters: Parameters that control how the model generates text, such as temperature and max tokens.
  • Temperature: A parameter that controls the randomness of the model's output.
  • Max Tokens: The maximum number of tokens that the model will generate.
  • API (Application Programming Interface): A set of rules and specifications that allow different software systems to communicate with each other.

7. Logical Connections:

  • The video starts by introducing VLM and its capabilities.
  • It then demonstrates three use cases: using Hugging Face models in code, serving Hugging Face models locally, and serving custom fine-tuned models.
  • Each use case builds upon the previous one, showing how VLM can be used for a variety of LLM deployment scenarios.

8. Data and Statistics:

  • The video mentions the size of the TinyLlama model (1.1 billion parameters).
  • It also mentions the use of 4-bit quantization for the fine-tuned model.

9. Conclusion:

VLM is a powerful and easy-to-use Python library for deploying and serving LLMs. It simplifies the process of using Hugging Face models and allows users to deploy their own custom fine-tuned models with minimal effort. By providing an OpenAI-compatible API, VLM enables seamless integration with existing tools and workflows. The key takeaways are the ease of use, the ability to serve both standard and fine-tuned models, and the OpenAI API compatibility.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.