Gemini 3 (FINAL Checkpoints Tested): I TESTED EVERY CHECKPOINT of Gemini-3. It's dropping this month

By AICodeKing

Share:

Key Concepts

  • Gemini 3.0 Pro: The upcoming iteration of Google's Gemini large language model, with early indications suggesting its release is imminent.
  • Checkpoints: Different versions or stages of a model's development and testing, often used in A/B testing and previews.
  • AI Studio: A platform for developing and testing AI models, including Gemini.
  • Vertex AI: Google Cloud's managed machine learning platform.
  • Spatial Reasoning: The ability of a model to understand and manipulate objects in three-dimensional space, crucial for tasks like floor plan generation.
  • One-shot Generation: The ability of a model to produce a complete and coherent output from a single prompt, without iterative refinement.
  • Tool Calling: The capability of a model to invoke external tools or functions to perform specific tasks.
  • Quantization: A technique used to reduce the precision of model weights, often leading to smaller model sizes and faster inference, but potentially impacting performance.
  • Consistency: The degree to which a model produces similar outputs for the same or similar prompts across multiple runs.
  • Latency: The time it takes for a model to generate a response, particularly the time to the first token.
  • Generative Code Polish: The quality and sophistication of code generated by a model.
  • Web OS Style Demos: Simple web-based demonstrations that are often easy for many frontier models to execute and not indicative of true reasoning capabilities.
  • Gemini CLI: A command-line interface for interacting with Gemini models.
  • Jules: Likely referring to a specific agent or framework that utilizes Gemini models.

Gemini 3.0 Pro: Early Indicators and Performance Observations

The video discusses the imminent release of Gemini 3.0 Pro, evidenced by a brief listing on Vertex AI. The author has tested various checkpoints, observing significant performance variations.

Early A/B Testing and the 2HT Checkpoint

  • Initial Discovery: Hints of Gemini 3.0 Pro emerged through A/B tests in AI Studio, where users occasionally received a different checkpoint (identified by a 2HT prefix in network logs) when selecting Gemini 2.5 Pro.
  • Rarity: This 2HT checkpoint appeared infrequently, approximately once every 40-50 prompts.
  • Impressive Performance: When accessed, the 2HT checkpoint demonstrated remarkable capabilities:
    • Spatial Reasoning: Generated sensible floor plans with correctly placed doors and furniture.
    • Creative Generation: Produced well-composed SVG images (panda with a burger) and 3D scenes (Pokéball in 3.js with excellent lighting).
    • Simulation: Showcased a high-quality Minecraft-like scene and a good butterfly simulation, though GPT-5 was noted as slightly better in the latter.
    • Reasoning: Excelled at AIM-style questions and riddles, exhibiting behavior suggestive of a "thinking variant" due to slower first-token generation.
  • Leaderboard Impact: This checkpoint achieved the top position on the author's leaderboard, showing a roughly 25% improvement over Sonnet 4.5. The author was willing to pay Sonnet pricing for this level of performance.

The "Weird Middle Chapter": The ECPT Checkpoint

  • Checkpoint Swapping: Google began swapping checkpoints, leading to a perceived decline in quality with an ECPT checkpoint.
  • Performance Degradation: This checkpoint felt "nerfed," exhibiting:
    • Lower quality floor plans.
    • Less cohesive SVG panda.
    • Functional but less intelligent chess gameplay (dumb captures).
    • Reduced polish in the 3.js Pokéball demo.
    • Laggy and flat Minecraft scene.
    • Mediocre butterfly simulation.
    • Loss of good lighting and camera setup in the Blender Pokéball.
  • Misleading Benchmarks: The author cautions against relying on simple web-style demos (HTML/CSS) as benchmarks, as these are easily handled by most frontier models. True differentiation lies in complex tasks like 3.js with real math, spatial logic, and deeper reasoning.
  • Hypothesis: The ECPT checkpoint was speculated to be a "Flash" or a low-thinking Pro variant, possibly quantized for broader rollout, safety testing, or latency optimization.
  • Comparison: While still better than Sonnet for many developer workflows, it did not feel like a generational upgrade.

The Bounceback: The X28 Checkpoint

  • Improved Performance: The X28 checkpoint marked a significant rebound, perceived as a step above 2HT.
  • Enhanced Capabilities: Re-testing with the same prompts and adding new ones revealed:
    • Superior Spatial Reasoning: Floor plans were described as "real," with proper doors, sensible layouts, and improved furniture with lighting controls. Consistency across runs was a notable improvement over Sonnet.
    • Visual Polish: The SVG panda appeared to be actively eating, and the 3.js Pokéball featured colorful backgrounds and better polish.
    • Advanced Simulations: The Minecraft scene included rivers and cleaner lighting, and the butterfly simulation was excellent with added details like rocks and flowers.
    • Functional Demos: A Rust CLI for image conversion worked, and the Blender script regained proper lighting and camera setup.
    • Complex UI/Agent Tasks: A degree of separation network simulation with sliders and regeneration controls was executed flawlessly with a clean UI, demonstrating attention to typography and spacing.
    • Tool Calling: Initial tool selection via RU's human relay was accurate, showing promise for agent capabilities.
  • Quantifiable Improvement: The author estimated a 5-10% improvement over 2HT and a significant leap over current Sonnet and other models on their "real prompts."

Quirks and Observations Across Checkpoints

  • "Thinking Variants": Stronger checkpoints exhibit characteristics of deliberate thought, indicated by a slow first token followed by steady output, even without visible traces.
  • Consistency: High-end checkpoints demonstrate unusually good consistency, which is crucial for app developers requiring near-deterministic behavior.
  • UI and Visual Taste: Models show strong design sensibilities, picking appropriate fonts and layouts, moving away from generic aesthetics.
  • Tool Calling as a Hinge: While raw reasoning is excellent, reliable chaining of function calls in live agents is critical for practical applications. Training for Gemini CLI and Jules patterns could make this a powerful pairing.
  • Purpose of Nerfed Checkpoints: The existence of ECPT is attributed to testing safety, latency, and serving limits. The hope is that the public release will be closer to X28 or 2HT.

Pricing Expectations and Market Positioning

  • Sonnet Pricing Justification: If Gemini 3 Pro lands around Sonnet pricing, the author believes the performance justifies it.
  • Higher Pricing Justification: If priced above Sonnet, Google must demonstrate reliability in tool calling, strong throughput, and consistent quality over long sessions.
  • Below Sonnet Pricing: Pricing below Sonnet would attract a significant user base, given the strength of the current product ecosystem (Gemini CLI, Jules, AI Studio generators).
  • Bottleneck Removal: Gemini 3 Pro's success hinges on its ability to remove the model as a bottleneck in the existing ecosystem.

Comparative Performance

  • Code Generation: Best Gemini 3 checkpoints are at or above Opus for generative code polish.
  • Spatial Reasoning & 3D: Clearly ahead of Sonnet 4.5.
  • Math & Consistency: Competitive with GPT-5.
  • Caveats: GPT-5 may edge out on certain physics-like simulations, but Gemini 3's consistency and UI taste are advantageous for practical tasks.
  • Launch Build Impact: The quality of the launch build (closer to X28/2HT vs. ECPT) will determine if it's a "new 3.5 Sonnet moment."

Release Timing and Benchmarking Advice

  • Imminent Preview: The Vertex listing with "11-2025" suggests an imminent preview.
  • Release Strategy: Google is likely to release a Pro preview first, followed by Flash, and then iterative updates. Ultra's release is uncertain.
  • Focus on Served Performance: The author prioritizes actual served performance (reasoning depth, tool call accuracy, consistency under load) over model labels.
  • Benchmarking Best Practices:
    • Avoid Web OS Demos: Do not rely on simple web-style demos for judging capability.
    • Push Complex Tasks: Test with 3D, math, and multifile tool flows.
    • Check Response Stability: Measure consistency across regenerations.
    • Measure Latency: Track both first-token latency and consistent handling of prompts after retries.
    • Agent Testing: Evaluate planning across steps, not just single function calls.

Conclusion and Future Outlook

The author expresses cautious optimism based on the observed performance of the stronger Gemini 3 Pro checkpoints (2HT and X28). If the public release mirrors these, it will be the leading mainstream model for developers. A release closer to the ECPT variant would still be good but not the significant leap anticipated. The author plans to conduct a full benchmark suite upon public preview release on Vertex AI, focusing on token economics, latency, and tool call pass rates to establish a clear price-to-performance picture.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video