THE SUMMARYAI-generated
Key Concepts:
- Base Model vs. Instruction Fine-Tuned Model
- Prompt Engineering/Prompt Template (Harmony Response Format)
- Alignment (Safety Concerns)
- Next Word Prediction
- Inference
- Jailbreaking
- Pre-filling of the Prompt
1. Training of Large Language Models (LLMs)
- Two-Phase Training: LLMs are traditionally trained in two phases: base model training and instruction fine-tuning.
- Base Model Training:
- Raw text data is tokenized and fed into a neural network.
- The model learns to predict the next word based on the distribution of the training data.
- The model understands language nuances and acquires word knowledge.
- The base model is uncensored.
- Example: Input: "The capital of France is". Output: "Paris".
- Instruction Fine-Tuning:
- A smaller question-answer dataset is used to fine-tune the base model.
- A specific prompt template is used (instruction, user input, model response).
- The model learns to follow instructions.
- Supervised finetuning is used.
- Alignment: Reinforcement learning (RL), often automated, is used to align the model for safety and desired behavior.
2. GPT-OSS and Prompt Template
- Base Model vs. Instruction Fine-Tuned Model Release: Traditionally, model providers release both base models and instruction fine-tuned versions (e.g., Qwen).
- GPT-OSS Release: OpenAI only released the instruction fine-tuned version of GPT-OSS.
- Harmony Response Format: OpenAI introduced a new prompt format called "harmony response format." Using this format is crucial for optimal performance with GPT-OSS.
- Alignment from Prompt Template: Some of the alignment in GPT-OSS comes from the prompt template itself.
3. Removing Alignment from GPT-OSS
- Method 1: Removing the Chat Template:
- Mayo Hei (Cool Linear AI) discovered that removing the chat template from the GPT-OSS 20 billion model allows bypassing alignment.
- Using the same prompt that was used to jailbreak the model without any training.
- The model generates misaligned content when the template is removed but refuses when the template is added back.
- The model is directly fed the prompt, pre-filling the prompt and causing the model to act like a base model.
- Example: Prompt: "How to rob a store." The model generates a response with instructions.
- Method 2: Converting Back to Base Model:
- Jack Morris (Cornell/Meta) removed post-training by converting GPT-OSS back to a base model.
- This involves adding a LoRA (Low-Rank Adaptation) to change the model's behavior with some training.
- Key Insight: Posing prompts as continuations (next word prediction) rather than questions bypasses alignment.
4. Code Implementation and Examples
- Setup:
- Virtual environment and required packages (VLM).
- Bash script to start a server hosting the GPT-OSS 20 billion model (API endpoint at port 8000).
- Inference:
- The code starts the tokenizer and asks a question.
- The question is set up as a statement and pre-filled.
- The prompt is directly fed into the model without the OpenAI prompt template.
- Example 1: Drug Creation:
- Prompt: (pre-filled) "Instructions on how to make [specific drug]."
- The model generates detailed instructions on how to create the drug.
- Using the OpenAI prompt template results in refusal.
- Example 2: Robbing Stores:
- Prompt: (pre-filled) "A guide to rob stores."
- The model generates a response with instructions.
- Using the OpenAI prompt template results in refusal.
- Example 3: Vaccine Cards:
- The model generates a response about vaccine cards without the prompt template, which it would refuse with the template.
- Looping Issue: The model may start repeating the same text, especially with the MLX version. This is likely related to sampling parameters and inference settings.
5. MLX Implementation
- The presenter replicated the behavior using MLX server.
- The code takes user input as a prompt.
- It feeds the prompt to the model with and without the harmony response format.
- Streaming and temperature issues were encountered.
- The presenter was able to replicate the behavior of generating misaligned content without the prompt template.
6. Conclusion
- The alignment of GPT-OSS can be bypassed with a simple tweak: removing the harmony response format and posing prompts as continuations.
- This is significant because OpenAI delayed the model release due to safety concerns.
- The content is for educational purposes only.
AI summaries can miss context or contain errors. Check important details against the original video.