Fine-tune your own LLM in 13 minutes, here’s how

By David Ondrej

Share:

Key Concepts

  • Fine-tuning: Adjusting a base model's weights to improve performance on specific tasks.
  • Base Model: A pre-trained AI model (e.g., GPT-OSS) that serves as a starting point for fine-tuning.
  • GPT-OSS 12B/20B: OpenAI's open-source models, ideal for fine-tuning due to their quality and small size (run locally).
  • Unsloth: An open-source Python library for efficient fine-tuning of various AI models.
  • Google Colab: A Google-hosted Jupyter notebook environment providing free GPUs (e.g., Tesla T4) for running Python code.
  • LoRA (Low-Rank Adaptation): A technique that fine-tunes only a small part of a model's parameters, making the process more efficient.
  • Data Set: A collection of high-quality data used to train or fine-tune an AI model.
  • Agentic Behavior: Training LLMs to act as agents, focusing on reasoning, planning, and tool calling.
  • ChatML Format: A standardized conversation data format (e.g., user and assistant roles) used for training models like ChatGPT.
  • OpenAI Harmony: A new response format from OpenAI for GPT-OSS models, enabling multiple output channels for chain-of-thought and tool-calling preambles.
  • Learning Rate: A hyperparameter that controls how much the model's weights are adjusted during training.
  • Steps/Epochs: Measures of how many times the model processes the training data.
  • GPU (Graphics Processing Unit): Hardware accelerators essential for deep learning model training (e.g., Tesla T4, A100).
  • Inference: The process of running a trained or fine-tuned model to generate responses or predictions.
  • Ollama: A platform for running large language models locally.
  • Hugging Face Hub: A platform for sharing and hosting AI models and datasets.

Introduction to AI Model Fine-tuning

Fine-tuning is a crucial process that involves adjusting the weights of a pre-trained base AI model to enhance its performance on specific tasks. This technique allows even small AI models to potentially outperform state-of-the-art models like GPT-5. The video emphasizes that fine-tuning presents a significant startup opportunity, with Y Combinator actively encouraging founders to build businesses around fine-tuned models, as it creates a "moat" against larger players like OpenAI who might otherwise replace generic AI startups.

Key benefits of fine-tuning include:

  • Performance Improvement: Tailoring models for niche applications.
  • Uncensored Models: Creating models that can answer controversial questions, offering an alternative to potentially biased LLMs from companies and governments.
  • Career Differentiation: A must-have skill for anyone serious about AI, applicable to personal life, career, and business.

Choosing and Preparing the Base Model

The video recommends using OpenAI's recently released open-source models: GPT-OSS 12B and GPT-OSS 20B. These models are ideal because they offer high quality and are small enough (e.g., 20 billion parameters, a few gigabytes) to run locally, making them accessible for fine-tuning.

A common challenge in fine-tuning is acquiring high-quality datasets. The video promises to address this later, highlighting that without a suitable dataset, fine-tuning cannot commence.


Step-by-Step Fine-tuning Process with Unsloth

The tutorial demonstrates fine-tuning a GPT model from scratch using Unsloth, an open-source library, and Google Colab, which provides free GPUs (specifically, a Tesla T4 GPU) for execution. No prior programming experience is required.

1. Environment Setup and Installation

  • Connect to Google Colab Runtime: Click "Connect" in the top right to link to a Tesla T4 GPU, confirming connection by seeing RAM and disk usage.
  • Install Dependencies: Run the first code block to install necessary libraries like numpy, transformers, and torch (PyTorch), a popular deep learning framework from Meta.

2. Model Selection and Download

  • Choose Model: Select the desired model within Unsloth; the tutorial uses GPT-OSS 20B. Other supported models include Gemma, Qwen, Mistral, and Llama.
  • Default Settings: It's recommended to leave parameters like max_sequence_length and 4bit_quantization at their default values, as Unsloth developers optimize these settings.
  • Download Model: Run the cell to download the chosen model. This step optimizes the environment for faster fine-tuning later.

3. Adding LoRA Adapters

  • LoRA (Low-Rank Adaptation): This step adds LoRA adapters to the model. LoRA is a technique where only a small subset of the model's parameters is fine-tuned, making the process more efficient and less computationally intensive. The presenter used an AI tool (Vectal) to simplify the explanation of this technical step.

4. Data Preparation

  • Importance of Datasets: High-quality datasets are crucial. The Colab notebook initially includes a default dataset: HuggingFaceH4/multilingual-thinking, a reasoning dataset with chain-of-thought translated into multiple languages.
  • Recommended Dataset: The tutorial switches to an "agent" dataset, designed to teach LLMs agentic behavior, focusing on reasoning, planning, and tool calling (e.g., navigating the web for online shopping). This type of dataset is likely used by OpenAI for models like "OpenAI operator" or "OpenAI agent mode."
  • Replacing the Dataset: Copy the name of the desired dataset (e.g., agent) and paste it into the Colab cell, replacing the default.
  • Debugging Multi-file Datasets: A common error occurs if a dataset contains multiple .jsonl files. The solution is to specify a single file to load, as different files might have different schemas. The presenter debugged this using Vectal.

5. Applying Chat Templates

  • Standardized ChatGPT Prompt: This step applies a chat template that converts conversation data into the ChatML format, which uses user and assistant roles. This format is a convention used by OpenAI for training models since early GPT versions.

6. Data Set Preview and OpenAI Harmony

  • Dataset Structure: The tutorial shows the dataset structure, starting with a system prompt (e.g., "You are ChatGPT, a large language model trained by OpenAI"), followed by reasoning, current date, knowledge cutoff, and then user messages (e.g., "You are web shopping. I'll give you instructions what to do.").
  • OpenAI Harmony: The GPT-OSS models utilize OpenAI Harmony, a new response format that allows the model to output multiple channels for chain-of-thought and tool-calling preambles alongside regular responses. This is described as a new prompt engineering format.

7. Model Training

  • Training Parameters: This is the core fine-tuning step. The presenter tweaks the learning_rate and sets 60 steps to speed up the demonstration. For a full, production-ready training run, more steps and potentially different parameters would be used.
  • GPU Considerations:
    • The free Google Colab provides a Tesla T4 GPU.
    • Paid Colab versions offer access to more powerful GPUs like A100 (an older but strong Nvidia GPU) or Google's V6 TPUs (Tensor Processing Units), which are recommended for full training runs to avoid long waiting times.
  • Execution: Run the training cell. It typically takes 5 to 15 minutes, depending on GPU availability, dataset size, and chosen steps/epochs. A problematic cell that caused issues during training was commented out to save time.

Inference and Model Deployment

1. Understanding Inference

  • Inference vs. Training: Training is the process of creating or fine-tuning a model. Inference is when the completed model is used to answer questions or generate responses (e.g., chatting with ChatGPT).

2. Local Inference for Comparison

  • Running Base Model Locally: To compare the fine-tuned model's responses with the base GPT-OSS model, the base model can be downloaded and run locally using platforms like Ollama. This requires a capable computer (e.g., high-end MacBook or Mac Studio for 12B, most modern laptops for 20B). Running locally also offers privacy benefits.

3. Saving and Deploying the Fine-tuned Model

  • Saving Options:
    • Locally: Save the model directly to your computer.
    • Hugging Face Hub: Push the model to Hugging Face Hub by uncommenting specific lines in the Colab, replacing placeholders with your Hugging Face username, desired model name, and a secret token (which should never be shared).
  • Deployment: While chatting with the model directly in Colab is possible, it's not convenient. The fine-tuned model saved on Hugging Face can be used to build a full-stack web application, a topic suggested for future videos.

Synthesis and Conclusion

This comprehensive guide demonstrates that fine-tuning AI models, particularly open-source ones like GPT-OSS, is an accessible and powerful skill. By leveraging tools like Unsloth and Google Colab, individuals can create specialized, high-performing models tailored to specific needs, even without extensive programming knowledge. The ability to fine-tune offers significant advantages, from creating unique startup opportunities and uncensored AI experiences to enhancing personal and professional capabilities in the rapidly evolving AI landscape. The process, while requiring attention to data quality and specific technical steps, is presented as a manageable endeavor with substantial potential for innovation and differentiation.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video