QWEN-3: EASIEST WAY TO FINE-TUNE WITH REASONING πŸ™Œ

Prompt EngineeringAbout 5 min readMay 6, 2025Watch original
THE SUMMARYAI-generated

Fine-Tuning Qwen3 Models: A Comprehensive Guide

Key Concepts:

  • Qwen3 models: Large language models with hybrid reasoning capabilities.
  • Fine-tuning: Adapting a pre-trained model to a specific dataset.
  • Catastrophic forgetting: The tendency of models to forget initial knowledge during continual fine-tuning.
  • LoRA (Low-Rank Adaptation): A parameter-efficient fine-tuning technique that adds adapter weights instead of modifying the original model weights.
  • Hybrid reasoning: The ability to enable or disable reasoning within the same model using hyperparameters.
  • Chain of Thought (CoT): A reasoning technique where the model generates intermediate steps to arrive at the final answer.
  • Quantization: Reducing the memory footprint of a model by using lower-precision data types.
  • Prompt engineering: Designing effective prompts to elicit desired responses from the model.

1. Introduction to Fine-Tuning Qwen3 Models

  • Qwen3 models are a good option for fine-tuning due to their different sizes, allowing them to run on smartphones or scale to large clusters.
  • Fine-tuning Qwen models on custom datasets requires a specific data structure due to their hybrid thinking mode.
  • The video covers setting up custom datasets, optimal inference hyperparameters, and key considerations for fine-tuning Qwen models.

2. Fine-Tuning with Unsloth

  • Unsloth is presented as a good option for fine-tuning LLMs.
  • The video uses an official notebook from the Unsloth team.
  • Unsloth reduces memory footprint by quantizing models using Dynamic 2.0 quantization.
  • Dynamic 2.0 quantization is supported within Llama.cpp, LlamaIndex, and Open Web.

3. Addressing Catastrophic Forgetting

  • The video references a paper on catastrophic forgetting in LLMs during continual fine-tuning.
  • The paper shows that fine-tuning a model on separate tasks can lead to forgetting initial knowledge.
  • LoRA is used to mitigate catastrophic forgetting by adding adapter weights instead of modifying the original weights.
  • LoRA adapters are smaller in dimension compared to the original model weights.

4. Code Walkthrough and Data Preparation

  • The most important parts of the codebase are data preparation and combining reasoning traces with non-reasoning data.
  • Combining reasoning and non-reasoning data is important to preserve reasoning capabilities.
  • If reasoning capabilities are not a concern, an instruct dataset can be used.

5. Installation and Model Loading

  • The video demonstrates installing Unsloth.
  • To run the Qwen3 14B model, the model name is provided (hosted on Hugging Face).
  • The maximum sequence length (context window) is defined. Qwen3 supports up to 128,000 tokens, but the example limits it to 2048 tokens.
  • The model is loaded in 4-bit, and LoRA adapters are used for fine-tuning.

6. Hyperparameter Tuning

  • The video briefly mentions hyperparameters like rank and Lora alpha.
  • Rank determines the size of the LoRA adapter matrices.
  • Lora alpha determines the impact of the LoRA adapters on the original weights.

7. Data Preparation: Reasoning and Non-Reasoning Datasets

  • Qwen3 models support hybrid reasoning, which can be turned on or off.
  • To preserve reasoning capabilities, the dataset should combine reasoning traces (chain of thought) and non-reasoning data.
  • The example combines a chain-of-thought dataset from R1 and an instruct dataset from ShareGPT.
  • The chain-of-thought dataset includes the original problem, chain-of-thought traces, and the final answer.
  • The non-reasoning dataset includes human input and responses from a chatbot.

8. Data Set Details

  • The reasoning data set contains fields like "expected answer", "model generation", and "original problem".
  • The non-reasoning data set contains "conversations", "source", and "scores".
  • A sample reasoning trace includes the original problem, chain-of-thought traces with special tags, and the final solution.

9. Prompt Templates for Qwen3

  • The video emphasizes the importance of understanding the prompt template of Qwen3.
  • The prompt template is specific to the post-trained model, not the pre-trained base model.
  • To enable thinking, the prompt needs to include specific thinking tags.
  • The video shows the prompt template for the fine-tune version of Qwen3.

10. Data Conversion and Standardization

  • The reasoning data is converted into the Qwen3 prompt template format.
  • The non-reasoning data is also standardized and converted into the same format.
  • The non-reasoning dataset is sampled to create a balanced dataset.
  • All examples are structured as a single string within a data frame with a key of "text".

11. Training the Model

  • The model is trained similarly to an instruct fine-tune version.
  • The model learns to produce thinking versus non-thinking tokens based on the prompt template.
  • The SFT trainer is set up with the model name, tokenizer, combined dataset, and hyperparameters.
  • The example uses a small batch size and runs for a maximum of 30 steps.
  • The training uses about 12 GB of VRAM, allowing fine-tuning on a free T4 GPU from Google Colab.

12. Inference and Hyperparameter Settings

  • The same tokenizer is used for inference.
  • Specific hyperparameters are recommended for thinking versus non-thinking mode.
  • For non-thinking mode, a higher temperature (0.7) is recommended.
  • For thinking mode, a smaller temperature is recommended.
  • The video emphasizes the importance of using the appropriate hyperparameters for optimal performance.

13. Inference Examples

  • The video demonstrates disabling thinking mode and generating a response without thinking tokens.
  • It then enables thinking mode and generates a response with thinking tokens.
  • The inference method is similar to the QWEN-VL model and Gemini 2.5 Flash.

14. Saving and Loading LoRA Adapters

  • The LoRA adapter can be saved using the save_pretrained function.
  • The tokenizer also needs to be saved.
  • The LoRA adapter can be loaded by providing its name.

15. Conclusion

  • The video concludes by highlighting the exciting possibilities of fine-tuning Qwen models for specific needs.
  • The 6 billion parameter model can be fine-tuned and run on edge devices like smartphones.
  • The Google Colab notebook link is provided in the video description.

Key Takeaways

  • Fine-tuning Qwen3 models requires careful data preparation and understanding of the prompt template.
  • LoRA is an effective technique for fine-tuning while mitigating catastrophic forgetting.
  • Hybrid reasoning capabilities allow for controlling thinking versus non-thinking modes using hyperparameters.
  • Unsloth provides a convenient way to fine-tune Qwen models with reduced memory footprint.
  • Fine-tuned Qwen models can be deployed on various devices, including edge devices.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.