Key Concepts
- GBD5 (GPT5): OpenAI's latest model, focusing on real-world utility and affordability.
- Chart Crimes: Inaccurate or misleading data visualizations in OpenAI's presentations.
- Agentic Coding: The ability of a model to autonomously perform coding tasks.
- Rate Limits: Restrictions on the number of messages users can send within a specific time frame.
- Unified Router: A system that automatically selects the appropriate model based on the complexity of the query.
- Arc AGI: A benchmark used to measure a model's general intelligence.
- OpenAI Pull Requests Benchmark: An internal benchmark measuring performance on pull request tasks.
- OpenAI Proof: An internal benchmark representing internal research and engineering bottlenecks.
- MLE Bench: An agentic benchmark measuring the model's capability to solve Kaggle computation problems.
- Sweet Lancer: A benchmark evaluating performance on real-world economically valuable full-stack engineering tasks.
- Paper Bench: A benchmark assessing a model's ability to replicate the results of a research paper.
- Healthbench: A benchmark focused on healthcare-related tasks.
- LMIS ELO Score: A metric used to evaluate the performance of language models.
- Intelligence Pareto Frontier: The boundary representing the optimal balance between intelligence and cost.
- Continuous Trained Real-Time Router Model: A model that continuously learns and adapts to route queries to the most appropriate model.
- Single-Shot Application Development: Building an application in a single attempt or interaction.
- Context Window: The amount of text a model can consider when generating a response.
- Tokens: Units of text used for processing by language models.
- 4-bit Floating-Point Precision: A method of representing numbers with reduced precision to lower computational costs.
- AGI: Artificial General Intelligence.
Data Visualization Issues ("Chart Crimes")
- The presentation of GBD5 contained several instances of misleading data visualization.
- Examples include bars of the same height representing different values (e.g., 30 and 69), and a bar representing 52 being taller than one representing 69.
- Another instance involved 50 being lower than 47 on a plot.
- The agentic coding capabilities plot had incorrect labeling, which was later corrected.
- The official blog post initially showed GPD Nano outperforming GPD5 on high settings, but this was later fixed.
- In the OpenAI proof benchmark, the green box has a higher height than the blue one, even though they are exactly the same 2%.
Access and Rate Limits
- OpenAI announced that GBD5 would be accessible to everyone for free, but rate limits apply.
- Plus users have a rate limit of 80 messages every 3 hours.
- Users explicitly using "GPD5 thinking" get 200 messages per week, the same as 03.
- Free tier users get only 10 messages every 5 hours.
- Pro and Teams users get "unlimited" access, but the definition of "unlimited" is not specified.
- Paid users on chat GPD have only two options: GPD5 or GPD5 thinking.
- The smart router, which is supposed to select a model based on complexity, is currently broken, routing queries to smaller models like GPT5 mini or nano.
Benchmarks and Performance
- Arc AGI: GBD5 achieves state-of-the-art performance, second only to Grock 4, at a substantially lower cost.
- OpenAI Pull Requests Benchmark: Performance is similar to OpenAI 03 when browsing is disabled.
- OpenAI Proof: Performance is similar to 03, despite being a real-world representation of internal bottlenecks.
- MLE Bench: Marginal improvement compared to 03.
- Sweet Lancer: Marginal improvement compared to 03; chart GPT agent performs better.
- Paper Bench: No significant improvements; performance degradation in some areas (e.g., capture the map).
- Healthbench: GPT5 thinking mini outperforms GP5 thinking.
Model Selection and Complexity
- The new GPD5 series simplifies model selection, but the explanation of the new models compared to previous ones is confusing for API users.
- The jump from GPT4 to GPT5 may not be as significant as previous jumps (e.g., GPT3 to GPT4) for most users.
- The real power of GPD5 is likely to be felt on very complex tasks, similar to Gemini 2.5 deep think.
Price to Performance Ratio
- GBD5 is a powerful model, especially for agentic coding capabilities.
- It is creating a new Pareto frontier for cost to performance ratio.
- GBD5 has left the Gemini series from Google behind in terms of price to performance.
- The unified router can automatically select models based on needs, but its effectiveness is yet to be determined.
- GBD5 is the best option available at the moment in terms of price to performance.
Technical Specifications and Pricing
- GBD5 has a 400,000 token context window with 128,000 max output tokens.
- It is multimodal in nature for input, but the output is only text.
- Pricing is $125 per input and $10 per output.
- Cash processing makes it even cheaper.
- GBD5 is similar to Gemini 2.5 Pro for less than 200,000 tokens.
- The mini or nano version is even cheaper and may offer better performance than models like flash.
OpenAI's Strategy and Future
- OpenAI is serving GBD5 models in 4-bit precision to reduce costs.
- Unifying everything into a single system saves serving and GPU costs.
- The release may be underwhelming for a version jump from 4 to 5.
- Sam Altman stated that the focus is on real-world utility and mass accessibility/affordability.
- OpenAI likely has better models but cannot host them at a reasonable price.
Testing and Community Feedback
- A classic farmer problem test revealed that the model struggles to take only the goat to the other side of the river.
- The model did draw a diagram, which is a neat feature.
- The presenter is interested in community feedback on whether GBD5 is better than Opus and the Cloud series, on par with them, or underwhelming.
Synthesis/Conclusion
GBD5 represents an incremental improvement over previous models, with a strong focus on affordability and real-world utility. While there are concerns about misleading data presentations and the effectiveness of the unified router, GBD5 excels in agentic coding and offers a compelling price-to-performance ratio. OpenAI's strategy appears to be centered on mass accessibility, even if it means holding back potentially more powerful models. The community's experience and feedback will be crucial in determining the true value and impact of GBD5.
AI summaries can miss context or contain errors. Check important details against the original video.





