Gemini 3 Pro (Fully Tested): This MODEL Broke MY BENCHMARKS! Better than X58 & BEST AI CODER YET!

By AICodeKing

Share:

Key Concepts

  • Gemini 3
  • Benchmarks (state-of-the-art, SWBench verified)
  • Pricing (input/output tokens, cost per token)
  • Availability (Google AI Studio, API, Gemini CLI)
  • Anti-gravity (agentic AI)
  • Kingbench questions (floor plan, SVG Panda, pokeball, chessboard, Minecraft clone, butterfly, CLI tool, Blender script, mathematics, riddles)
  • Agentic benchmark
  • Kilo Code
  • ZenMox
  • Nano Banana

Gemini 3 Launch and Performance

Gemini 3 has been launched and is now available on Google's AI Studio. The model is described as "state-of-the-art" across almost all benchmarks, outperforming models like Sonnet and GPT 4.1. While it didn't excel on SWBench verified, the speaker dismisses this benchmark as "rigged."

Benchmarks and Performance

  • General Performance: Gemini 3 is reported to be state-of-the-art on most benchmarks, surpassing Sonnet and GPT 4.1.
  • SWBench Verified: The model did not perform as well on SWBench verified, which the speaker considers unreliable.
  • Speaker's Benchmark Results: The speaker conducted their own benchmark tests, focusing initially on "Kingbench questions" rather than agentic ones. The agentic benchmark video is planned for release later.

Pricing and Availability

  • Pricing Structure:
    • Under 200,000 token limit: $2 for input, $12 for output.
    • Over 200,000 token limit: $4 for input, $18 for output.
  • Cost-Effectiveness: The pricing is considered "quite good" and cheaper than Sonnet, offering excellent price-to-performance.
  • Availability:
    • Google AI Studio
    • API
    • Gemini CLI (for free usage)

New Feature: Anti-gravity

A new feature called "anti-gravity" has been launched, which appears to be Google's proprietary agentic AI. A separate video will be dedicated to this feature.

Gemini 3 Performance on Kingbench Questions

The speaker details Gemini 3's performance on various "Kingbench" questions, highlighting its capabilities:

  • Floor Plan Question: The model generated a good floor plan with correctly laid out rooms. It also demonstrated the ability to change the time of day, with lights turning on at night, which was described as "insane." While not perfect, it was close to X58 and superior in presentation.
  • SVG Panda Holding a Burger: The output was "so good," surpassing even X58, and looked like a realistic picture, with detailed rendering of the burger.
  • Pokeball in 3JS: Described as "insane" and "perfect," with a very good visual representation.
  • Chessboard with Autoplay: The model created a chessboard that resembled boards from chess.com, featuring "topnotch" autoplay and animations. The fact that it was generated in "one shot" was considered "unbelievable."
  • Minecraft Clone in Kandinsky Style: The model "nailed this," producing a good environment with "insane" movement.
  • Majestic Butterfly Flying in the Garden: This was considered "even better than the X58 checkpoint," with realistic butterfly behavior, good aesthetics, fine camera movement, and well-rendered flowers.
  • CLI Tool for Converting Images in Rust: The model generated a "really good" CLI tool, demonstrating great knowledge and speed. The speaker suggested OpenAI could learn from this.
  • Blender Script: The model successfully created a Blender scene, aligning the camera, adding lights, and even generating the "Pokémon capturing waves thing." This was deemed "insane."
  • Mathematics Questions: Gemini 3 passed both mathematics questions on the first try, and also solved a riddle, achieving a "100%" score on the speaker's benchmark.

Benchmark Conclusion and Future Plans

  • Kingbench 2.0 Retirement: With Gemini 3's performance, Kingbench 2.0 is officially retired.
  • Agentic Benchmark: Results from the agentic benchmark will be shared soon after testing is complete.
  • New Benchmarks: The speaker plans to share new benchmarks in the next video.
  • Overall Satisfaction: The speaker expresses high satisfaction with the Gemini 3 release, stating it's the "true end of Claude GPT 4.1 and blah blah models." They hope other companies will focus on creating good models rather than inflating prices and prioritizing profit. Google has set a high standard.

How to Use Gemini 3

  • Coder Integration:
    • Gemini CLI: Directly usable.
    • Kilo Code: Works "out of the box" after installation. The speaker uses it here and notes it comes with $10 of free credit.
  • Free API Access:
    • ZenMox: Has added Gemini 3 for free with rate limits. The speaker has not tested this API personally but confirms its availability.

Other Mentions and Future Content

  • Nano Banana: This model is not yet launched but is expected soon.
  • Anti-gravity: The speaker will test this agentic AI and share results. They are seeking viewer input on whether to release the agent testing video or the anti-gravity video first.
  • Anti-gravity Functionality: It appears to be a basic fork with AI features, capable of navigating the browser to perform tasks. The speaker needs more time to form a definitive opinion.

Conclusion and Call to Action

The speaker reiterates that Gemini 3 is a "really good" model with "a ton of power." They encourage viewers to stay tuned for more benchmarking and the anti-gravity editor video. Viewers are asked to share their thoughts in the comments, subscribe to the channel, and consider donating via Super Thanks or joining as a channel member for perks.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video