THE SUMMARYAI-generated
Key Concepts
- Fine-tuning: Further training a pre-trained large language model (LLM) on a specific dataset to improve its performance on a particular task.
- Unsloth: A Python library for accelerated fine-tuning of LLMs, reducing VRAM usage and speeding up training.
- VRAM: Video RAM, the memory on a GPU.
- QLoRA: Quantized Low-Rank Adapters, a memory-efficient fine-tuning technique.
- Gradient Checkpointing: A technique to reduce VRAM usage during training by recomputing activations.
- Ollama: A tool for running LLMs locally.
- Hugging Face Datasets: A library for easily accessing and processing datasets for machine learning.
- SFT Trainer: Supervised Fine-Tuning Trainer from the TRL (Transformer Reinforcement Learning) library.
- LoRA: Low-Rank Adaptation, a parameter-efficient fine-tuning technique that uses low-rank matrices to update the model.
- GGUF: A file format for storing quantized LLMs, compatible with llama.cpp and Ollama.
Fine-Tuning LLMs Locally with Unsloth
Introduction to Unsloth and Fine-Tuning
- The video demonstrates how to fine-tune large language models (LLMs) locally using the Unsloth library in Python.
- Unsloth accelerates fine-tuning, reduces VRAM usage by 50-80% through techniques like 4-bit quantization, gradient checkpointing, and CPU offloading.
- Fine-tuning involves taking a pre-trained model (e.g., Fi, Mistral, Qwen) and further training it on a specific dataset.
- LoRA and QLoRA techniques are used to tune the model parameters via low-rank matrices.
- Fine-tuning can improve performance on specific tasks but may decrease performance on general tasks.
Hardware and Software Requirements
- Fine-tuning requires a GPU with sufficient VRAM and CUDA support.
- A 3060Ti with 8GB of VRAM is sufficient for the demonstration.
- Google Colab can be used for free if local hardware is insufficient, using a T4 GPU runtime.
- Required Python packages:
unsloth,datasets,trl,jupyterlab. - A virtual environment is recommended (using
venvoruv).
Choosing a Base Model and Dataset
- Select an Unsloth-compatible model from the Unsloth collections on Hugging Face.
- The video uses
unsloth/fi3-mini-4k-instruct-bnb-4bit. - The dataset should be structured with "prompt" and "response" fields.
- The example dataset (
people_data.json) contains prompts with descriptive text about a person and responses with JSON objects containing the person's name, age, job, and gender. - The dataset can be custom-generated or sourced from existing datasets.
Setting Up the Environment
- In Google Colab, change the runtime type to a GPU.
- Install necessary packages using
pip install unslothoruv add unsloth. - Install Jupyter Lab for interactive coding.
Loading and Formatting the Data
- Import
jsonanddatasetsfrom thehuggingfacelibrary. - Load the JSON data from the
people_data.jsonfile. - Format the data into strings with user and assistant tags:
<|user|>\n{prompt}\n<|assistant|>\n{json.dumps(response)}\n<|file_separator|>
- Create a Hugging Face dataset from the formatted strings using
Dataset.from_dict. - Professional Approach: Use the chat template provided by the tokenizer instead of custom tags for better compatibility and to avoid problematic tokens.
Loading the Model and Tokenizer
- Import
FastLanguageModelfromunsloth. - Load the pre-trained model and tokenizer using
FastLanguageModel.from_pretrained:model_name: The model identifier (e.g.,unsloth/fi3-mini-4k-instruct-bnb-4bit).max_sequence_length: Maximum sequence length (e.g., 2048).data_type: Set toNonefor automatic selection.load_in_4bit: Set toTrueto load the 4-bit quantized model.
Preparing the Model for Fine-Tuning with LoRA
- Prepare the model for parameter-efficient fine-tuning using
FastLanguageModel.get_peft_model. - Specify LoRA parameters:
r: Rank of the LoRA matrices (e.g., 64).target_modules: Modules where LoRA matrices are injected (e.g.,["query_key_value"]).lora_alpha: Scaling factor (e.g.,r * 2).lora_dropout: Dropout probability (e.g., 0.0).bias: Set toNoneto apply only to weights.use_gradient_checkpointing: Set toTrueto use Unsloth's gradient checkpointing.
Training the Model
- Import
SFTTrainerandSFTConfigfromtrl. - Create an
SFTTrainerinstance:model: The prepared model.train_dataset: The Hugging Face dataset.tokenizer: The tokenizer.dataset_text_field: The name of the text field in the dataset (e.g., "text").max_seq_length: The maximum sequence length.sft_config: AnSFTConfiginstance with training parameters:per_device_train_batch_size: Batch size per GPU (e.g., 2).gradient_accumulation_steps: Number of steps to accumulate gradients (e.g., 4).warmup_steps: Number of warmup steps (e.g., 10).max_steps: Maximum training steps (e.g., 60).num_train_epochs: Number of training epochs (e.g., 3).logging_steps: Logging frequency (e.g., 1).output_dir: Output directory for logs and checkpoints (e.g., "outputs").optim: Optimizer (e.g., "adamw_bnb_8bit").
- Start the training process using
trainer.train().
Using the Fine-Tuned Model in Python
- Set the model to inference mode using
FastLanguageModel.for_inference_model(model). - Craft an example message as a list of dictionaries with "role" and "content" keys.
- Apply the chat template using
tokenizer.apply_chat_template. - Generate the model's output using
model.generate:input_ids: The tokenized input.max_new_tokens: Maximum number of tokens to generate (e.g., 512).use_cache: Set toTrueto use caching.temperature: Sampling temperature (e.g., 0.7).do_sample: Set toTrueto enable sampling.top_p: Top-p sampling value (e.g., 0.9).
- Decode the output using
tokenizer.batch_decodeand print the response.
Exporting the Model for Ollama
- Save the model in GGUF format using
model.save_pretrained_gguf:save_directory: The directory to save the model (e.g., "fine-tuned-model").tokenizer: The tokenizer.quantization_method: Quantization method (e.g., "Q4_K_M").max_memory_usage: Limit GPU memory usage (e.g., 0.3).
Loading and Running the Model in Ollama
- Install Ollama from the official website.
- Create a
Modelfilewith the following content:FROM fine-tuned-model/unsloth.q4_k_m.gguf
- Create the Ollama model using
ollama create fine-tuned-model -f Modelfile. - Run the model using
ollama run fine-tuned-model. - Chat with the model in the command line.
Conclusion
- The video demonstrates the process of fine-tuning LLMs locally using Unsloth and deploying them with Ollama.
- Fine-tuning can improve performance on specific tasks, but requires careful selection of the base model, dataset, and training parameters.
- The process involves loading and formatting data, preparing the model for fine-tuning, training the model, and exporting it for deployment.
- The fine-tuned model can be used in Python or deployed locally with Ollama.
- Proper fine-tuning is a science and an art, requiring experimentation and optimization.
AI summaries can miss context or contain errors. Check important details against the original video.





