OpenAI Finally Goes Open-Source: 120B & 20B Models

Prompt EngineeringAbout 5 min readAug 6, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Open weight models, 120 billion parameter model, 20 billion parameter model, agentic tasks, coding tool usage, reasoning models, chain of thought, Olama, Apache 2.0 license, H100 GPU, consumer grade hardware, floating point precision, mixture of experts, sparse models, context length, rotary positional embeddings, tokenizer, supervised finetuning, RL stage, safety standards, red teaming challenge, harmony prompt format, M2 Max, local GPT, Grock, open router.

Model Overview and Specifications

OpenAI has released two open weight models: a 120 billion parameter model and a 20 billion parameter model. The smaller model can run on consumer-grade hardware (laptop/desktop), while the larger model requires a high-end GPU like an H100 (80GB VRAM). The 20 billion model requires 16GB VRAM. Both models are released under the Apache 2.0 license and are designed for agentic tasks, coding tool usage, and reasoning. Users can control the level of reasoning effort (low, medium, high) via the system message.

  • Model Sizes: 120 billion and 20 billion parameters.
  • Licensing: Apache 2.0.
  • Intended Use: Agentic tasks, coding, reasoning.
  • Reasoning Control: Low, medium, or high reasoning efforts configurable via system message.
  • Accessibility: Available on Olama, Llama CPP, LM Studio, and through various API providers.
  • Hardware Requirements: 80GB VRAM (120B), 16GB VRAM (20B).
  • Floating Point Precision: 4-bit.

Training and Architecture

The models were trained using a similar setup to OpenAI's proprietary models. They are mixture of experts models, meaning only a fraction of the parameters are active per token (5 billion for the 120B model, 3.6 billion for the 20B model). This sparsity allows for a smaller memory footprint. The models support a context length of 128,000 tokens and use rotary positional embeddings. The 120 billion model has 128 experts, while the 20 billion model has 32 experts. Training data was primarily English text focused on STEM, coding, and general knowledge. The models use the same tokenizer as GPT-4o and GPT-4, called O 200K harmony. Fine-tuning involved supervised finetuning and an RL stage with high compute.

  • Architecture: Mixture of Experts (MoE).
  • Active Parameters: 5B (120B model), 3.6B (20B model) per token.
  • Context Length: 128,000 tokens.
  • Positional Embeddings: Rotary positional embeddings.
  • Training Data: Primarily English text (STEM, coding, general knowledge).
  • Tokenizer: O 200K harmony (same as GPT-4o and GPT-4).
  • Fine-tuning: Supervised finetuning and RL stage.

Performance and Benchmarks

The 120 billion parameter model's performance is close to GPT-4o mini on several key areas, especially with tool usage. Benchmarks include code-related tasks, math computation, and PhD-level science questions (MMLU). Test time scaling works, meaning increasing reasoning effort and token generation leads to consistent performance increases. The models' agentic capabilities are highlighted, with the 120 billion model performing similarly to GPT-4o mini.

  • Performance: 120B model close to GPT-4o mini.
  • Benchmarks: Code, math, MMLU.
  • Test Time Scaling: Increased reasoning effort and token generation improves performance.
  • Agentic Capabilities: 120B model similar to GPT-4o mini.

Safety and Openness

OpenAI emphasizes safety, hosting a red teaming challenge with a $500,000 prize fund to identify novel safety issues. The models' chain of thought is kept open without direct supervision to allow for monitoring of misbehavior, deception, and misuse. The release was coordinated with major players in the open-source space, ensuring day-one support on platforms like Olama, Llama CPP, and LM Studio.

  • Safety Focus: Red teaming challenge with $500,000 prize.
  • Chain of Thought: Intentionally kept open for monitoring misbehavior.
  • Community Collaboration: Coordinated release with major open-source platforms.

Practical Usage and Availability

The models are available on Hugging Face and can be run locally. OpenAI has released cookbooks for fine-tuning and handling the raw chain of thought. The models are natively quantized in MX 4-bit floating point precision. Implementations for running inference with PyTorch and on Apple Metal platform are also available. Partners include Hugging Face, VLM, Olama, Llama CVP, LM Studio, AWS, Fireworks, Together AI, Basten, B10, Data Bricks, Versel, and Cloudfire. Hardware optimization is ensured through collaboration with Nvidia, AMD, Cerebrris, and Croc.

  • Availability: Hugging Face, Olama, Llama CPP, LM Studio, etc.
  • Resources: Cookbooks for fine-tuning and chain of thought.
  • Quantization: MX 4-bit floating point precision.
  • Hardware Optimization: Collaborations with Nvidia, AMD, Cerebrris, Croc.

Initial Testing and Impressions

The 20 billion model was tested on an M2 Max with 96 GB of VRAM. The model identified itself as a conversational AI built on OpenAI's GPT-4 architecture. Grock open router reports 1200 tokens per second for the 20 billion model and 500 tokens per second for the 120 billion model. A concern is raised about the performance of the 4-bit floating point precision in actual coding tasks, which will be tested in a future video.

  • Test Environment: M2 Max with 96 GB VRAM.
  • Model Identification: Identified as GPT-4 based conversational AI.
  • Token Generation Speed (Grock): 1200 tokens/s (20B), 500 tokens/s (120B).
  • Concern: Coding performance with 4-bit precision.

Significance and Future Outlook

The release of these open weight models marks a significant step forward, providing advancements in reasoning capabilities and safety. Open models complement hosted models, offering a wider range of tools for research, innovation, and safer AI development. They lower the barrier for emerging markets and resource-constrained organizations. The release is seen as a response to open releases from Chinese companies like DeepSeek and Quinn. The future is predicted to be bright for open weight models, with hopes for more releases from Frontier Labs like Google and Anthropic.

  • Significance: Advances in reasoning and safety, complements hosted models.
  • Impact: Lowers barrier for emerging markets and resource-constrained organizations.
  • Market Context: Response to open releases from Chinese companies.
  • Future Outlook: Expectation of more open weight releases from other labs.

Conclusion

OpenAI's release of the 120 billion and 20 billion parameter open weight models represents a significant contribution to the open-source AI community. These models, designed for agentic tasks, coding, and reasoning, offer impressive performance comparable to GPT-4o mini, while also prioritizing safety through red teaming efforts and transparent chain of thought. The models' accessibility, combined with comprehensive documentation and broad platform support, promises to accelerate research, innovation, and the development of safer and more transparent AI systems.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.