GPT-OSS 120B + KingBench 2.0 (Tested): Worst of 2025? This Model is pretty bad at almost anything.

AICodeKingAbout 6 min readAug 6, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Open weights model
  • Mixture of Experts (MoE)
  • 4-bit quantization
  • Reasoning effort (low, medium, high)
  • Tool use
  • Chain of Thought (CoT)
  • Benchmarks
  • Ollama, LM Studio, VLLM
  • GLM 4.5 Air
  • 3JS, SVG, CLI
  • Open Router

Model Configuration and Capabilities

OpenAI has released two open weights models: a 117B parameter model and a 20B parameter model. Both are Mixture of Experts (MoE) models employing 4-bit quantization. The 117B model has 5.1 billion active parameters, while the 20B model has 3.6 billion active parameters. The 117B model can fit on a single 80GB GPU, and the 20B model on a 16GB GPU. These models are designed for reasoning tasks and support adjustable reasoning effort levels (low, medium, high). They excel at instruction following and tool use, providing access to the model's reasoning process for debugging. The models are released under the Apache license and support tool calling.

Accessibility and Usage

The models can be used with tools like Ollama, LM Studio, and VLLM. The GPTOSS 120B model performs comparably to OpenAI's GPT-4 Mini on reasoning benchmarks while running on a single 80GB GPU. The GPTOSS 20B model achieves similar results to OpenAI's GPT-3 Mini and can run on edge devices with 16GB of memory. A free demo is available on the GPToss site, requiring a Hugging Face account login. The models can be run locally using Ollama with a single command.

Provider Availability and Pricing

Several providers, including Gro, Cerebras, and Fireworks, offer access to these models. Gro charges approximately $15.75, while Cerebras charges around $25.09. The models are also accessible via Open Router, although reasoning effort cannot be adjusted through these providers yet.

Benchmark Performance

The presenter tested the models using a revamped benchmark suite with 10 new questions. Gemini 2.5 Pro achieved the highest score at 50%. The GPTOSS model scored low in the presenter's benchmarks, solving only one question (a riddle). The benchmark includes math problems (from Amy), coding challenges (primarily in 3JS), and SVG generation tasks. The model struggled with 3JS tasks, failing to render floor plans, SVGs of a panda with a burger, Pokeballs, chessboards with autoplay, web versions of Minecraft, and flying butterflies in a garden. It also failed to create a CLI tool in Rust for image conversion.

Comparison with GLM 4.5 Air

The presenter compared GPTOSS to GLM 4.5 Air, another open model with similar parameters. GLM 4.5 Air performed better overall, even though it scored only slightly higher in the benchmarks. For example, GLM 4.5 Air's floor plan generation was close to perfect, while GPTOSS failed to render anything. GLM 4.5 Air also performed better on the Pokeball, chessboard, Minecraft, and butterfly tests. The CLI tool for image conversion also worked in GLM 4.5 Air.

Key Arguments and Perspectives

The presenter argues that while OpenAI's open weights model is a positive step, its performance is not as impressive as existing open models like GLM 4.5 Air. GLM 4.5 Air can run on the same hardware as GPTOSS while delivering better results. The presenter plans to evaluate both models with reasoning enabled in the future. The initial impression of GPTOSS is not favorable, but the smaller model might perform better.

Notable Quotes

  • "[The GPTOSS 120B model] achieves near par with OpenAI 04 Mini on core reasoning benchmarks while running efficiently on a single 80GB GPU."
  • "[The GPTOSS 20B model] delivers similar results to OpenAI 03 Mini on common benchmarks and can run on edge devices with just 16GB of memory making it ideal for ondevice use cases, local inference or rapid iteration without costly infrastructure."
  • "Even without reasoning, evaluating both tells you that though OpenAI's model is good, it is not anywhere near what open models that we already have."

Technical Terms and Concepts

  • Open Weights Model: A machine learning model whose weights (parameters) are publicly available, allowing for modification, redistribution, and use without licensing restrictions.
  • Mixture of Experts (MoE): A neural network architecture that combines multiple sub-networks (experts) to handle different parts of the input space, improving performance and scalability.
  • 4-bit Quantization: A technique to reduce the memory footprint and computational cost of a model by representing its weights using only 4 bits per parameter, leading to compression and faster inference.
  • Reasoning Effort: The amount of computational resources and time allocated to the model's reasoning process, which can be adjusted to balance accuracy and latency.
  • Chain of Thought (CoT): A prompting technique that encourages the model to generate intermediate reasoning steps before providing the final answer, improving its ability to solve complex problems.
  • Tool Use: The ability of a language model to interact with external tools or APIs to gather information or perform actions, extending its capabilities beyond text generation.
  • Benchmarks: Standardized tests used to evaluate the performance of machine learning models on specific tasks, allowing for comparison and progress tracking.
  • Ollama, LM Studio, VLLM: Tools and frameworks for running and deploying large language models locally or in production environments.
  • GLM 4.5 Air: An open-source language model developed by Zhipu AI, known for its strong performance and efficiency.
  • 3JS: A JavaScript library for creating and displaying animated 3D computer graphics in a web browser.
  • SVG: Scalable Vector Graphics, an XML-based vector image format for defining graphics in a web browser.
  • CLI: Command-Line Interface, a text-based interface for interacting with a computer operating system or application.
  • Open Router: A platform that provides a unified API for accessing multiple language models from different providers.

Logical Connections

The video begins by introducing OpenAI's new open weights model and its configurations. It then discusses the model's accessibility and usage, followed by a comparison with other models, particularly GLM 4.5 Air. The presenter uses benchmark results to support the argument that while OpenAI's model is a step forward, it does not outperform existing open models in several key areas. The video concludes with the presenter's overall impression and future plans for further evaluation.

Synthesis/Conclusion

OpenAI's release of an open weights model is a significant development, offering potential for wider accessibility and customization. However, the presenter's benchmarks and comparisons with GLM 4.5 Air suggest that the model's performance, particularly in coding and reasoning tasks, may not yet match that of other available open models. While the model's ease of use and integration with tools like Ollama are positive aspects, further improvements are needed to achieve state-of-the-art performance. The presenter plans to continue evaluating the model with reasoning enabled to gain a more comprehensive understanding of its capabilities.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.