Gemini 3 (FINAL Checkpoints Tested): I TESTED EVERY CHECKPOINT of Gemini-3. It's dropping this month
By AICodeKing
Key Concepts
- Gemini 3.0 Pro: The upcoming iteration of Google's Gemini large language model, with early indications suggesting its release is imminent.
- Checkpoints: Different versions or stages of a model's development and testing, often used in A/B testing and previews.
- AI Studio: A platform for developing and testing AI models, including Gemini.
- Vertex AI: Google Cloud's managed machine learning platform.
- Spatial Reasoning: The ability of a model to understand and manipulate objects in three-dimensional space, crucial for tasks like floor plan generation.
- One-shot Generation: The ability of a model to produce a complete and coherent output from a single prompt, without iterative refinement.
- Tool Calling: The capability of a model to invoke external tools or functions to perform specific tasks.
- Quantization: A technique used to reduce the precision of model weights, often leading to smaller model sizes and faster inference, but potentially impacting performance.
- Consistency: The degree to which a model produces similar outputs for the same or similar prompts across multiple runs.
- Latency: The time it takes for a model to generate a response, particularly the time to the first token.
- Generative Code Polish: The quality and sophistication of code generated by a model.
- Web OS Style Demos: Simple web-based demonstrations that are often easy for many frontier models to execute and not indicative of true reasoning capabilities.
- Gemini CLI: A command-line interface for interacting with Gemini models.
- Jules: Likely referring to a specific agent or framework that utilizes Gemini models.
Gemini 3.0 Pro: Early Indicators and Performance Observations
The video discusses the imminent release of Gemini 3.0 Pro, evidenced by a brief listing on Vertex AI. The author has tested various checkpoints, observing significant performance variations.
Early A/B Testing and the 2HT Checkpoint
- Initial Discovery: Hints of Gemini 3.0 Pro emerged through A/B tests in AI Studio, where users occasionally received a different checkpoint (identified by a
2HTprefix in network logs) when selecting Gemini 2.5 Pro. - Rarity: This
2HTcheckpoint appeared infrequently, approximately once every 40-50 prompts. - Impressive Performance: When accessed, the
2HTcheckpoint demonstrated remarkable capabilities:- Spatial Reasoning: Generated sensible floor plans with correctly placed doors and furniture.
- Creative Generation: Produced well-composed SVG images (panda with a burger) and 3D scenes (Pokéball in 3.js with excellent lighting).
- Simulation: Showcased a high-quality Minecraft-like scene and a good butterfly simulation, though GPT-5 was noted as slightly better in the latter.
- Reasoning: Excelled at AIM-style questions and riddles, exhibiting behavior suggestive of a "thinking variant" due to slower first-token generation.
- Leaderboard Impact: This checkpoint achieved the top position on the author's leaderboard, showing a roughly 25% improvement over Sonnet 4.5. The author was willing to pay Sonnet pricing for this level of performance.
The "Weird Middle Chapter": The ECPT Checkpoint
- Checkpoint Swapping: Google began swapping checkpoints, leading to a perceived decline in quality with an
ECPTcheckpoint. - Performance Degradation: This checkpoint felt "nerfed," exhibiting:
- Lower quality floor plans.
- Less cohesive SVG panda.
- Functional but less intelligent chess gameplay (dumb captures).
- Reduced polish in the 3.js Pokéball demo.
- Laggy and flat Minecraft scene.
- Mediocre butterfly simulation.
- Loss of good lighting and camera setup in the Blender Pokéball.
- Misleading Benchmarks: The author cautions against relying on simple web-style demos (HTML/CSS) as benchmarks, as these are easily handled by most frontier models. True differentiation lies in complex tasks like 3.js with real math, spatial logic, and deeper reasoning.
- Hypothesis: The
ECPTcheckpoint was speculated to be a "Flash" or a low-thinking Pro variant, possibly quantized for broader rollout, safety testing, or latency optimization. - Comparison: While still better than Sonnet for many developer workflows, it did not feel like a generational upgrade.
The Bounceback: The X28 Checkpoint
- Improved Performance: The
X28checkpoint marked a significant rebound, perceived as a step above2HT. - Enhanced Capabilities: Re-testing with the same prompts and adding new ones revealed:
- Superior Spatial Reasoning: Floor plans were described as "real," with proper doors, sensible layouts, and improved furniture with lighting controls. Consistency across runs was a notable improvement over Sonnet.
- Visual Polish: The SVG panda appeared to be actively eating, and the 3.js Pokéball featured colorful backgrounds and better polish.
- Advanced Simulations: The Minecraft scene included rivers and cleaner lighting, and the butterfly simulation was excellent with added details like rocks and flowers.
- Functional Demos: A Rust CLI for image conversion worked, and the Blender script regained proper lighting and camera setup.
- Complex UI/Agent Tasks: A degree of separation network simulation with sliders and regeneration controls was executed flawlessly with a clean UI, demonstrating attention to typography and spacing.
- Tool Calling: Initial tool selection via RU's human relay was accurate, showing promise for agent capabilities.
- Quantifiable Improvement: The author estimated a 5-10% improvement over
2HTand a significant leap over current Sonnet and other models on their "real prompts."
Quirks and Observations Across Checkpoints
- "Thinking Variants": Stronger checkpoints exhibit characteristics of deliberate thought, indicated by a slow first token followed by steady output, even without visible traces.
- Consistency: High-end checkpoints demonstrate unusually good consistency, which is crucial for app developers requiring near-deterministic behavior.
- UI and Visual Taste: Models show strong design sensibilities, picking appropriate fonts and layouts, moving away from generic aesthetics.
- Tool Calling as a Hinge: While raw reasoning is excellent, reliable chaining of function calls in live agents is critical for practical applications. Training for Gemini CLI and Jules patterns could make this a powerful pairing.
- Purpose of Nerfed Checkpoints: The existence of
ECPTis attributed to testing safety, latency, and serving limits. The hope is that the public release will be closer toX28or2HT.
Pricing Expectations and Market Positioning
- Sonnet Pricing Justification: If Gemini 3 Pro lands around Sonnet pricing, the author believes the performance justifies it.
- Higher Pricing Justification: If priced above Sonnet, Google must demonstrate reliability in tool calling, strong throughput, and consistent quality over long sessions.
- Below Sonnet Pricing: Pricing below Sonnet would attract a significant user base, given the strength of the current product ecosystem (Gemini CLI, Jules, AI Studio generators).
- Bottleneck Removal: Gemini 3 Pro's success hinges on its ability to remove the model as a bottleneck in the existing ecosystem.
Comparative Performance
- Code Generation: Best Gemini 3 checkpoints are at or above Opus for generative code polish.
- Spatial Reasoning & 3D: Clearly ahead of Sonnet 4.5.
- Math & Consistency: Competitive with GPT-5.
- Caveats: GPT-5 may edge out on certain physics-like simulations, but Gemini 3's consistency and UI taste are advantageous for practical tasks.
- Launch Build Impact: The quality of the launch build (closer to
X28/2HTvs.ECPT) will determine if it's a "new 3.5 Sonnet moment."
Release Timing and Benchmarking Advice
- Imminent Preview: The Vertex listing with "11-2025" suggests an imminent preview.
- Release Strategy: Google is likely to release a Pro preview first, followed by Flash, and then iterative updates. Ultra's release is uncertain.
- Focus on Served Performance: The author prioritizes actual served performance (reasoning depth, tool call accuracy, consistency under load) over model labels.
- Benchmarking Best Practices:
- Avoid Web OS Demos: Do not rely on simple web-style demos for judging capability.
- Push Complex Tasks: Test with 3D, math, and multifile tool flows.
- Check Response Stability: Measure consistency across regenerations.
- Measure Latency: Track both first-token latency and consistent handling of prompts after retries.
- Agent Testing: Evaluate planning across steps, not just single function calls.
Conclusion and Future Outlook
The author expresses cautious optimism based on the observed performance of the stronger Gemini 3 Pro checkpoints (2HT and X28). If the public release mirrors these, it will be the leading mainstream model for developers. A release closer to the ECPT variant would still be good but not the significant leap anticipated. The author plans to conduct a full benchmark suite upon public preview release on Vertex AI, focusing on token economics, latency, and tool call pass rates to establish a clear price-to-performance picture.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

GPT 5.6 SOL: TBH, IT'S OKAY.. I have SERIOUS CONCERNS.
AICodeKing

GLM 5.2 Is INSANE. Better than Claude Fable 5?
Zubair Trabzada | AI Workshop

GLM-5.2 (Fully Tested): I got EARLY ACCESS & This MODEL is CRAZY!
AICodeKing

20 days of compute vs 7 hours: rethinking what state-of-the-art means — Bertrand Charpentier, Pruna
AI Engineer

API vs Subscriptions vs Local: I Measured Intelligence Per Dollar.
Eduards Ruzga

Google Just Dropped COSMO Then Mysteriously Pulled It
AI Revolution

Gemini 3.5 Flash In Arena! POWERFUL, Cheap, & Fast NEW AI Model! (Fully Tested)
WorldofAI