Fine-Tuning Local LLMs with Unsloth & Ollama

NeuralNineAbout 5 min readAug 24, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Fine-tuning: Further training a pre-trained large language model (LLM) on a specific dataset to improve its performance on a particular task.
  • Unsloth: A Python library for accelerated fine-tuning of LLMs, reducing VRAM usage and speeding up training.
  • VRAM: Video RAM, the memory on a GPU.
  • QLoRA: Quantized Low-Rank Adapters, a memory-efficient fine-tuning technique.
  • Gradient Checkpointing: A technique to reduce VRAM usage during training by recomputing activations.
  • Ollama: A tool for running LLMs locally.
  • Hugging Face Datasets: A library for easily accessing and processing datasets for machine learning.
  • SFT Trainer: Supervised Fine-Tuning Trainer from the TRL (Transformer Reinforcement Learning) library.
  • LoRA: Low-Rank Adaptation, a parameter-efficient fine-tuning technique that uses low-rank matrices to update the model.
  • GGUF: A file format for storing quantized LLMs, compatible with llama.cpp and Ollama.

Fine-Tuning LLMs Locally with Unsloth

Introduction to Unsloth and Fine-Tuning

  • The video demonstrates how to fine-tune large language models (LLMs) locally using the Unsloth library in Python.
  • Unsloth accelerates fine-tuning, reduces VRAM usage by 50-80% through techniques like 4-bit quantization, gradient checkpointing, and CPU offloading.
  • Fine-tuning involves taking a pre-trained model (e.g., Fi, Mistral, Qwen) and further training it on a specific dataset.
  • LoRA and QLoRA techniques are used to tune the model parameters via low-rank matrices.
  • Fine-tuning can improve performance on specific tasks but may decrease performance on general tasks.

Hardware and Software Requirements

  • Fine-tuning requires a GPU with sufficient VRAM and CUDA support.
  • A 3060Ti with 8GB of VRAM is sufficient for the demonstration.
  • Google Colab can be used for free if local hardware is insufficient, using a T4 GPU runtime.
  • Required Python packages: unsloth, datasets, trl, jupyterlab.
  • A virtual environment is recommended (using venv or uv).

Choosing a Base Model and Dataset

  • Select an Unsloth-compatible model from the Unsloth collections on Hugging Face.
  • The video uses unsloth/fi3-mini-4k-instruct-bnb-4bit.
  • The dataset should be structured with "prompt" and "response" fields.
  • The example dataset (people_data.json) contains prompts with descriptive text about a person and responses with JSON objects containing the person's name, age, job, and gender.
  • The dataset can be custom-generated or sourced from existing datasets.

Setting Up the Environment

  • In Google Colab, change the runtime type to a GPU.
  • Install necessary packages using pip install unsloth or uv add unsloth.
  • Install Jupyter Lab for interactive coding.

Loading and Formatting the Data

  • Import json and datasets from the huggingface library.
  • Load the JSON data from the people_data.json file.
  • Format the data into strings with user and assistant tags:
    • <|user|>\n{prompt}\n<|assistant|>\n{json.dumps(response)}\n<|file_separator|>
  • Create a Hugging Face dataset from the formatted strings using Dataset.from_dict.
  • Professional Approach: Use the chat template provided by the tokenizer instead of custom tags for better compatibility and to avoid problematic tokens.

Loading the Model and Tokenizer

  • Import FastLanguageModel from unsloth.
  • Load the pre-trained model and tokenizer using FastLanguageModel.from_pretrained:
    • model_name: The model identifier (e.g., unsloth/fi3-mini-4k-instruct-bnb-4bit).
    • max_sequence_length: Maximum sequence length (e.g., 2048).
    • data_type: Set to None for automatic selection.
    • load_in_4bit: Set to True to load the 4-bit quantized model.

Preparing the Model for Fine-Tuning with LoRA

  • Prepare the model for parameter-efficient fine-tuning using FastLanguageModel.get_peft_model.
  • Specify LoRA parameters:
    • r: Rank of the LoRA matrices (e.g., 64).
    • target_modules: Modules where LoRA matrices are injected (e.g., ["query_key_value"]).
    • lora_alpha: Scaling factor (e.g., r * 2).
    • lora_dropout: Dropout probability (e.g., 0.0).
    • bias: Set to None to apply only to weights.
    • use_gradient_checkpointing: Set to True to use Unsloth's gradient checkpointing.

Training the Model

  • Import SFTTrainer and SFTConfig from trl.
  • Create an SFTTrainer instance:
    • model: The prepared model.
    • train_dataset: The Hugging Face dataset.
    • tokenizer: The tokenizer.
    • dataset_text_field: The name of the text field in the dataset (e.g., "text").
    • max_seq_length: The maximum sequence length.
    • sft_config: An SFTConfig instance with training parameters:
      • per_device_train_batch_size: Batch size per GPU (e.g., 2).
      • gradient_accumulation_steps: Number of steps to accumulate gradients (e.g., 4).
      • warmup_steps: Number of warmup steps (e.g., 10).
      • max_steps: Maximum training steps (e.g., 60).
      • num_train_epochs: Number of training epochs (e.g., 3).
      • logging_steps: Logging frequency (e.g., 1).
      • output_dir: Output directory for logs and checkpoints (e.g., "outputs").
      • optim: Optimizer (e.g., "adamw_bnb_8bit").
  • Start the training process using trainer.train().

Using the Fine-Tuned Model in Python

  • Set the model to inference mode using FastLanguageModel.for_inference_model(model).
  • Craft an example message as a list of dictionaries with "role" and "content" keys.
  • Apply the chat template using tokenizer.apply_chat_template.
  • Generate the model's output using model.generate:
    • input_ids: The tokenized input.
    • max_new_tokens: Maximum number of tokens to generate (e.g., 512).
    • use_cache: Set to True to use caching.
    • temperature: Sampling temperature (e.g., 0.7).
    • do_sample: Set to True to enable sampling.
    • top_p: Top-p sampling value (e.g., 0.9).
  • Decode the output using tokenizer.batch_decode and print the response.

Exporting the Model for Ollama

  • Save the model in GGUF format using model.save_pretrained_gguf:
    • save_directory: The directory to save the model (e.g., "fine-tuned-model").
    • tokenizer: The tokenizer.
    • quantization_method: Quantization method (e.g., "Q4_K_M").
    • max_memory_usage: Limit GPU memory usage (e.g., 0.3).

Loading and Running the Model in Ollama

  • Install Ollama from the official website.
  • Create a Modelfile with the following content:
    • FROM fine-tuned-model/unsloth.q4_k_m.gguf
  • Create the Ollama model using ollama create fine-tuned-model -f Modelfile.
  • Run the model using ollama run fine-tuned-model.
  • Chat with the model in the command line.

Conclusion

  • The video demonstrates the process of fine-tuning LLMs locally using Unsloth and deploying them with Ollama.
  • Fine-tuning can improve performance on specific tasks, but requires careful selection of the base model, dataset, and training parameters.
  • The process involves loading and formatting data, preparing the model for fine-tuning, training the model, and exporting it for deployment.
  • The fine-tuned model can be used in Python or deployed locally with Ollama.
  • Proper fine-tuning is a science and an art, requiring experimentation and optimization.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.