Gemini 3.0 Pro (ECPT Checkpoint - TESTED) : They NERFED IT! 10% LOW SCORE but still good.
By AICodeKing
Key Concepts ECPT checkpoint, Gemini 3.0 Pro, Zenith checkpoint (GPT5), Sonnet (Claude 3 Sonnet), Kilo code, 3JS (JavaScript 3D library), Web OS, Flexible divs, Python interpreter, Easter egg snake game, Volutrics, Blender script, Quantized model, Lowinking/reasoning variant, Flash (model variant), VO 3.1.
Introduction: Disappointment with Google's ECPT Checkpoint
The speaker expresses significant disappointment and skepticism regarding Google's latest ECPT checkpoint, a new version of Gemini 3.0 Pro. Despite considerable hype, the speaker finds this checkpoint to be "nerfed" or a "lower thinking variant" compared to previous checkpoints they have tested. This performance decline evokes "flashbacks of the Zenith checkpoint of GPT5," which was superior to GPT5 but never publicly released. The speaker was compelled to test this checkpoint due to numerous requests but found its performance underwhelming and, in some cases, buggy.
Critique of Performance Across Benchmarks
The speaker evaluated the ECPT checkpoint using a series of benchmarks, highlighting specific deficiencies:
- Floor Plan Generation: The model's output was deemed "very mediocre, if not bad, and it's very bland." Rooms were "not correctly aligned" and fell short of the quality observed in previous checkpoints.
- SVG Panda Rendering: This version was "not like the previous version," appearing less "good or polished." While the burger element was "pretty good," the surrounding panda was not.
- 3JS Pokéball: Although "pretty great" and similar to previous outputs, it was "not as great as the previous checkpoints," with the background and other elements lacking in quality.
- Chess Game Logic: The chessboard functionality was "not very great either." While it worked, "most of the moves here are pretty dumb," failing to capture pieces effectively and performing worse than prior checkpoints.
- 3JS Minecraft Game: This application was "not extraordinary." It functioned but was "a bit laggy," lacked sufficient lighting, and had "volutrics [that] are not great."
- Butterfly in Garden Animation: The output was described as "fine," neither "very good" nor "very bad," but again, "not as great as the previous checkpoints."
- Blender Script for Pokéball: This script was "okay," but unlike previous versions, it lacked lighting. Despite this, the "dimensions and everything are good."
- General Questions (Pentagon): The model successfully "passes all these general questions," which is considered "pretty great." However, the speaker notes it seems "similar but worse," possibly due to quantization or being a lower-reasoning variant.
Debunking the "Web OS" Hype
A significant portion of the critique is directed at the "web OS" generation capabilities, which the speaker dismisses as a "gimmick" and a misleading benchmark for model capabilities.
- Comparison with Sonnet: The speaker demonstrates that Sonnet (presumably Claude 3 Sonnet) can generate a "web OS" in "one shot" using "Kilo code," producing results "as good as Gemini's generation being scattered around."
- Ease of "Web OS" Prompts: The speaker argues that prompts for generating a web OS are "relatively easy" and are not included in their standard benchmarks because they primarily involve "just elements with flexible divs and nothing more," rather than complex "mathematics for placing objects in 3D space" as required by 3JS questions.
- Misleading Hype from "Non-Programmers": The speaker attributes the excessive hype to "most of the dumb non-programmers" who post "dumb and silly things like that," failing to understand what constitutes a rigorous test of a model's capabilities.
- Not Exclusive to Gemini 3: The speaker explicitly states that generating a web OS in one prompt is "not something exclusive to Gemini 3" and can be achieved by "GPT5, Claude, and almost every other model."
Hypotheses for Nerfed Performance
The speaker speculates on several reasons for the perceived downgrade in the ECPT checkpoint's performance:
- Quantization: The model might have been "quantized to be deployed to general audiences," which often involves optimizing for efficiency at the cost of some performance.
- Safety Settings: The implementation of new safety settings could be limiting its capabilities.
- Lowinking/Reasoning Variant: It might be a less capable "lowinking or reasoning variant" of the full model.
- "Flash" Variant: The speaker suggests it "might be Flash," referring to a potentially lighter or faster, but less powerful, version of the model (e.g., Gemini Flash).
Overall Assessment and Comparison
Despite the criticisms, the speaker acknowledges that the model is "still really good" and "better than Sonnet" in some aspects. However, it "doesn't seem to be the new 3.5 sonnet moment anymore," indicating it's not a groundbreaking leap. The primary concern is that it is "a bit nerfed compared to the previous checkpoints," leading to skepticism. The model also exhibited "buggy" behavior, such as providing a "non-existing link" when asked to generate a floor plan.
Concerns and Future Outlook
The speaker expresses a general distrust in recent model launches due to "unnecessary hype" and models not living up to expectations. They hope that Google will release "not nerfed variants as well to try out" and fear that this launch might mirror the "GPT5 one," where a promising checkpoint never fully materialized or was downgraded. The speaker notes that "VO 3.1 is launching today" and "other models will also come pretty soon," indicating an active period for AI model releases.
Conclusion
The ECPT checkpoint of Gemini 3.0 Pro, while still a capable model, is perceived by the speaker as a "nerfed" version compared to earlier, more impressive checkpoints. The widespread hype, particularly around its ability to generate a "web OS," is deemed misleading, as such tasks are relatively simple and achievable by many advanced models. The speaker advises against judging a model's true capabilities based on "silly prompts" and emphasizes the need for more rigorous benchmarks, especially those involving complex 3D mathematics. The overall sentiment is one of disappointment and skepticism regarding the direction of recent AI model releases, with a hope for more robust and less compromised variants in the future.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

How the hometown humiliation of Putin marks a turning point for Ukraine | DW News
DW News

Shocking video shows moment paramedics are hit by Israel in 'double-tap' strike
Sky News

Putin Xi, To Catch a Castro, Red Carpet Rebellion • FRANCE 24 English
FRANCE 24 English

Trump's supporters furious over Trump smartphone scam.
ABC News In-depth

Nvidia Crushes Earnings again — What Jensen Huang sees next for AI
CGTN America

Samsung union suspends strike after reaching tentative pay deal • FRANCE 24 English
FRANCE 24 English

OH SH*T! The Banks are Dumping AI Loans!
Steven Van Metre