GPT-5.1 (Caterpillar CKPT Tested): This NEW GPT-5 Checkpoint by OpenAI seems QUITE GOOD!
By AICodeKing
Key Concepts
- Alleged OpenAI GPT 5.1 Models: New models purportedly from OpenAI, operating under stealth names.
- Stealth Names: Undisclosed or disguised names used for models, similar to Google's practice.
- Reasoning Budgets: A measure of a model's computational capacity for reasoning tasks, with higher budgets indicating greater potential.
- Design Arena & LM Arena: Platforms where users can interact with and test various AI models.
- Benchmarks: Standardized tests used to evaluate the performance of AI models across different tasks.
- GPT5 Codec: A specific OpenAI model previously used for in-depth planning and debugging.
- Model Degradation: A perceived decline in the performance of existing AI models over time.
- Quantization: A technique to reduce the size and computational requirements of AI models for inference.
- GPOSS Model: An OpenAI model that the speaker finds to be a marketing gimmick.
- Architecture and Passion: Factors the speaker believes are crucial for modern AI model development, beyond just training capacity.
- Gemini 3: Google's advanced AI model, presented as a benchmark for comparison.
Alleged OpenAI GPT 5.1 Models: Firefly, Chrysalis, Cicada, and Caterpillar
The video discusses the emergence of new, alleged OpenAI GPT 5.1 models operating under stealth names. These models are identified as Firefly, Chrysalis, Cicada, and Caterpillar. The core concept behind these models is that they are variations of the same underlying model, differentiated by their "reasoning budgets."
- Firefly: The model with the lowest reasoning budget.
- Chrysalis: Reported to have a reasoning juice of approximately 16.
- Cicada: Said to have a reasoning juice of around 64.
- Caterpillar: The model with the highest reasoning juice, at 256.
The speaker primarily tested the Caterpillar model, suggesting it is likely a new checkpoint for GPT5 or GPT5 Mini, pointing towards GPT 5.1.
Access and Usage
These alleged GPT 5.1 models are currently accessible through platforms like Design Arena and LM Arena. Design Arena allows users to input prompts and receive responses from one of the four models. While LM Arena also features these models, their appearance there is less frequent.
Benchmark Performance Analysis
The speaker provides a detailed analysis of the models' performance across various benchmarks, often comparing them to Gemini 3 and other models like Minimax, Claude, and GLM.
- Floor Plan Generation: The models perform poorly, unable to create useful floor plans.
- SVG Panda Eating a Burger: The output is described as "fine" but not comparable to Gemini 3.
- Pokeball in 3JS: The generation is acceptable but exhibits several issues.
- Chessboard: The model can generate a chessboard but struggles with keeping up with best moves, falling short of Gemini 3's capabilities.
- 3D Minecraft: This task fails entirely.
- Butterfly Flying in Garden: The generation is good but not novel, with Minimax producing a better butterfly.
- CLI Tool in Rust: The script works reasonably well, though with minor issues.
- Blender Script for Pokeball: The script does not work.
- Mathematics (Positive Integers, Convex Pentagon): The models excel at these tasks, providing correct answers.
- Riddle Solving: The models successfully solve riddles.
Overall Benchmark Assessment: The models are considered "fine" and better than Minimax and GLM, but perform slightly worse than Claude. They are significantly outperformed by Gemini 3 checkpoints. The speaker expresses disappointment, hoping these are "mini" models and not indicative of the full GPT5's capabilities, which also underperformed on their benchmarks.
Personal Use Cases and Model Degradation Concerns
The speaker previously relied on GPT5 Codec for in-depth planning and debugging, finding it superior to Opus. GPT5 Codec was effective for identifying and fixing issues when other models like Minimax, GLM, or Sonnet were stuck. It was also used for security checks. However, the speaker notes that GPT5 Codec has recently degraded in performance.
This degradation leads to the suspicion that OpenAI, like other model providers, might be intentionally making their models worse. This could be a strategic move to make newer models appear more impressive or a consequence of quantizing previous models to free up GPUs for new deployments. The lack of communication regarding these changes is criticized as a poor practice.
Limitations and Future Outlook
The speaker reiterates that for their personal use, GPT5 is only suitable for planning and not for other tasks. They acknowledge that Design Arena's system prompts might influence generations, but the consistent good performance on math questions suggests a genuine capability rather than just prompt engineering.
The speaker expresses a desire for OpenAI to succeed but feels they are "fumbling badly." The new nonprofit structure and the GPOSS model are viewed as marketing gimmicks, with the speaker preferring GLM4.5 air.
Shifting Paradigms in AI Development
A key argument presented is that modern AI model development is shifting from sheer training capacity to architecture and passion. This is exemplified by smaller companies like Minimax and ZAI pushing boundaries with capable, smaller models, contrasting with OpenAI's perceived reliance on "gimmicks."
The speaker expresses strong belief in Google due to their straightforward approach to AI, avoiding marketing gimmicks. Gemini's live mode is considered better than OpenAI's counterpart, and Google is seen as building a robust ecosystem around its models.
Conclusion and Call to Action
The speaker concludes that while the alleged GPT 5.1 models are "pretty cool" to test, their performance is not exceptional. They encourage viewers to test the models on Design Arena and share their experiences. The video ends with a call to subscribe, donate, and join the channel for perks.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

GPT 5.6 SOL: TBH, IT'S OKAY.. I have SERIOUS CONCERNS.
AICodeKing

GLM 5.2 Is INSANE. Better than Claude Fable 5?
Zubair Trabzada | AI Workshop

GLM-5.2 (Fully Tested): I got EARLY ACCESS & This MODEL is CRAZY!
AICodeKing

20 days of compute vs 7 hours: rethinking what state-of-the-art means — Bertrand Charpentier, Pruna
AI Engineer

API vs Subscriptions vs Local: I Measured Intelligence Per Dollar.
Eduards Ruzga

Google Just Dropped COSMO Then Mysteriously Pulled It
AI Revolution

Gemini 3.5 Flash In Arena! POWERFUL, Cheap, & Fast NEW AI Model! (Fully Tested)
WorldofAI