How to Fine Tune your own LLM using LoRA (on a CUSTOM dataset!)

Nicholas RenotteAbout 6 min readJun 11, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Fine-tuning: Adapting a pre-trained large language model (LLM) to a specific task or dataset.
  • LoRA (Low Rank Adaptation): A parameter-efficient fine-tuning technique that adds small, trainable matrices to the original LLM, rather than updating all the parameters.
  • Quantization: Reducing the precision of the model's weights (e.g., from 32-bit to 4-bit) to decrease memory usage and increase inference speed.
  • Synthetic Data Generation: Creating training data using an LLM to augment or replace real-world data.
  • Instruction Tuning: Training an LLM to follow instructions, typically using a dataset of question-answer pairs.
  • Chat Template: A specific format for structuring conversations with an LLM, including system prompts, user inputs, and assistant responses.
  • Greedy Decoding: A decoding strategy where the model always selects the most probable next token, leading to less creative but more predictable outputs.
  • RunPod: A cloud platform for renting GPU instances for machine learning tasks.
  • Ollama: A tool for running LLMs locally on your machine.
  • Langflow: A visual programming tool for building LLM-powered applications.

1. Introduction to Fine-Tuning and LoRA

  • Problem: Large LLMs are expensive to use and may not perform well on niche topics.
  • Solution: Fine-tuning allows you to adapt an LLM to a specific use case.
  • LoRA: A cost-effective fine-tuning method that adds "knowledge cartridges" to the LLM.
  • Quantization: Enables running fine-tuned LLMs on laptops.
  • Goal: To provide an end-to-end walkthrough of fine-tuning an LLM using LoRA on a custom dataset.

2. LoRA Explained

  • Full Fine-tuning: Updates all the weights in the LLM, which is computationally expensive.
  • LoRA: Introduces two smaller weight matrices (A and B) for each weight matrix to be tuned.
  • Rank: A hyperparameter in LoRA that determines the size of the adapter matrices. Higher rank = more parameters to update.
  • Matrix Multiplication: The matrices A and B are multiplied (BA) to match the dimensions of the original pre-trained weights.
  • Weight Addition: The resulting matrix BA is added to the original pre-trained weights.
  • Cost Consideration: Setting the rank too high can make LoRA more costly than full fine-tuning.

3. Baseline LLM Performance Evaluation

  • LLM Used: Llama 3.21B (small, fast, and cheap to fine-tune).
  • Tools: Langflow and O Lama.
  • Decoding Strategy: Top K = 1 (greedy decoding).
  • Test Questions (TM1-related):
    • What is a pick list in TM1?
    • What is the minimum number of dimensions for a cube in TM1?
    • How do I create a dimension in TM1?
    • What are the core advantages of TM1?
  • Results: The LLM exhibited hallucinations and inaccuracies, demonstrating the need for fine-tuning.

4. Synthetic Data Generation with Dockling and LightLLM

  • Data Source: A PDF document containing information about TM1.
  • Tool 1: Dockling: Used to load and chunk the PDF document.
    • DocumentConverter: Loads the PDF.
    • HybridChunker: Segments the document into smaller chunks.
    • contextualize method: Appends subject matter to the top of each chunk.
  • Tool 2: LightLLM: Used to generate synthetic question-answer pairs from the chunks.
  • LLM Used for Generation: Quen 2.514b (larger LLM).
  • Prompt Template: Used to instruct the LLM to generate question-answer pairs.
  • Data Structure: The generated data is structured as a list of dictionaries, each containing a question and an answer.
  • JSON Schema: Used to ensure that the LLM returns data in the correct format.

5. Data Pre-processing

  • Problem: The initial data format was not suitable for training.
  • Solution: A pre-processing script was created to clean and reformat the data.
  • Steps:
    • Loop through each chunk in the JSON data.
    • Extract the question and answer pairs.
    • Append the pairs to a new list.
    • Dump the list to a new JSON file (instruction.json).
  • Context Appending (Later Removed): Initially, the context of each chunk was appended to the question-answer pair. This was later found to decrease performance.

6. Setting up a GPU Instance on RunPod

  • Reason: Training requires a GPU.
  • Platform: RunPod (cloud GPU rental).
  • Instance Type: RTX A5000 (25 GB VRAM).
  • Connection: SSH connected over TCP.
  • Environment Setup:
    • Copy the data directory to the RunPod instance.
    • Install UV (package manager).
    • Initialize a new UV project.

7. Data Mapping and Batching for Training

  • Dependencies: datasets, transformers, torch, bitsandpl, colorama.
  • Loading Data: load_dataset from Hugging Face datasets.
  • Chat Template Formatting:
    • Define a system prompt.
    • Apply the Llama chat template to the tokenizer.
    • Loop through each sample in the batch.
    • Convert the question and answer into a JSON format with roles (system, user, assistant).
    • Apply the chat template to the JSON.
  • Tokenization: AutoTokenizer.from_pretrained is used to load the tokenizer.
  • Data Mapping: The map method is used to apply the format_chat_template function to the dataset.
  • Batching: The dataset is batched for efficient training.

8. Kicking off Training

  • Model Loading: AutoModelForCausalLM.from_pretrained is used to load the model.
  • Device Placement: The model is placed on the GPU (CUDA).
  • Quantization Configuration:
    • BitsAndBytesConfig is used to configure quantization.
    • load_in_4bit=True
    • double_quantization=True
    • quantization_type="nf4"
    • compute_dtype=torch.bfloat16
  • Gradient Checkpointing: Enabled to save memory.
  • LoRA Configuration:
    • LoraConfig is used to configure LoRA.
    • r=32 (rank)
    • lora_alpha=64
    • lora_dropout=0.05
    • target_modules=["all linear"]
    • task_type="CAUSAL_LM"
  • Trainer Setup:
    • SupervisedFineTuningTrainer is used for training.
    • Training arguments are specified using SFTConfig.
  • Training Execution: trainer.train() is called to start the training process.

9. Deploying Locally Using O Lama

  • Model Deployment: The fine-tuned model is deployed locally using O Lama.
  • Model File: A model file is created to specify the base model and the adapter file path.
  • O Lama Commands:
    • oama create <model_name> -f <model_file>: Creates a new O Lama model.
    • oama run <model_name>: Runs the O Lama model.
    • oama list: Lists the available O Lama models.
  • Langflow Integration: The O Lama model is integrated into Langflow for testing.

10. Improving Performance

  • Initial Performance: The initial fine-tuning resulted in some improvements, but the model still exhibited hallucinations.
  • Performance Improvement Strategies:
    1. Increasing Rank and Alpha: Increasing the rank and alpha of the LoRA adapter significantly improved performance.
    2. Adding More Data: Adding more data helped, but the quality of the data was crucial.
    3. Ultra-Targeted Data: Having less generic data but ultra-targeted data performed well.
    4. Data Quality Improvement: Using a data classification script to filter out low-quality data.
  • Final Training Parameters:
    • rank=256
    • lora_alpha=512
    • epochs=100
    • save_steps=1000
  • Data Quality Filtering:
    • A data quality script was created to classify each instruction example in terms of accuracy and style.
    • The script used an LLM to assign a score (1-10) for accuracy and style, with explanations.
    • Instruction pairs with scores below a threshold (e.g., 6 for both accuracy and style) were discarded.
  • Data Classification Script:
    • Used light_llm to classify instruction tuning records.
    • Classified between 1-10 in terms of accuracy and style.
    • Provided explanations for each score.
    • Penalized harmful, unhelpful, or dishonest instruction pairs.
  • Final Data Set: The final data set consisted of high-quality instruction pairs without context.

11. Conclusion

Fine-tuning large language models using LoRA can be a cost-effective way to adapt them to specific tasks. However, achieving good performance requires careful attention to data quality, training parameters, and deployment strategies. The key takeaways are:

  • LoRA is a parameter-efficient fine-tuning technique.
  • Quantization enables running fine-tuned LLMs on consumer hardware.
  • Synthetic data generation can augment or replace real-world data.
  • Data quality is crucial for achieving good performance.
  • Increasing the rank and alpha of the LoRA adapter can improve performance.
  • Data classification scripts can be used to filter out low-quality data.
  • O Lama and Langflow are useful tools for deploying and testing fine-tuned LLMs.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.