Gemini 3.0 Flash (Tested): Google's NEW Model is INTERESTING...

AICodeKingAbout 4 min readDec 26, 2025Watch original
THE SUMMARYAI-generated

Gemini 3.0 Flash: Benchmarks and Initial Observations

Key Concepts:

  • Gemini 3.0 Flash: A new, faster, and more cost-effective model in the Gemini 3 family, positioned as a lighter version of Gemini 3 Pro.
  • Multimodal Capabilities: The ability of the model to process and generate content from various input types (text, images, audio).
  • Auto Reasoning: A feature where the model automatically attempts to reason through a problem before providing an answer.
  • Agentic Benchmarks: Tests designed to evaluate a model’s ability to utilize tools effectively and make sensible decisions in a multi-step process.
  • Tool Calling: The model’s ability to identify and utilize external tools to assist in completing a task.
  • Input/Output Tokens: Units of text used to measure the cost of using a language model; input tokens are the text provided to the model, and output tokens are the text generated by the model.

Introduction of Gemini 3.0 Flash

Google has recently launched Gemini 3.0 Flash, currently available on platforms like Zenmucks (an open router). While official blog posts were pending at the time of the video, the model is positioned as a faster and more affordable alternative to Gemini 3 Pro. It’s designed for applications prioritizing speed and cost-efficiency, with a pricing structure of $0.30 per million input tokens and $0.50 per million output tokens. The model retains the core multimodal and reasoning capabilities of Gemini 3 Pro, built on the same architecture, but prioritizes responsiveness.

Technical Specifications & Design

Gemini 3 Flash is described as a “low latency model” optimized for “fast high throughput inference.” It supports native multimodal inputs – text, images, and audio – and incorporates the improved reasoning and long context handling features introduced with the Gemini 3 generation. It features “always reasoning” similar to Gemini 3 Pro, with adjustable reasoning budgets (high, etc.) though it defaults to optimal reasoning. The speaker emphasizes its strength in multimodal tasks, aligning with Gemini’s established leadership in visual and multimodal processing.

Benchmark Results: Non-Agentic General Questions

The speaker conducted initial benchmarks using a non-agentic, general question benchmark. The results demonstrate a mixed performance:

  • Floor Plan Generation: Gemini 3 Flash produced a poorly generated floor plan lacking essential features like doors and logical room arrangements, significantly inferior to Gemini 3 Pro.
  • SVG Panda with Burger: The model excelled at generating an SVG image of a panda with a burger, achieving results comparable to Gemini 3 Pro in terms of detail and overlapping elements. However, the speaker still preferred the Gemini 3 Pro version for its more lifelike rendering with shadows.
  • Pokeball in 3JS: Gemini 3 Flash outperformed Gemini 3 Pro in generating a Pokeball in 3JS, accurately representing the black stripes.
  • Chessboard with Autoplay: The model failed to generate a functional chessboard with autoplay functionality, even with multiple attempts. Gemini 3 Pro successfully completed this task.
  • Minecraft Clone in 3JS: Similar to the chessboard, the Minecraft clone generation was unsuccessful.
  • Majestic Butterfly: The butterfly generation was considered “kind of good,” with a visually appealing butterfly but limited, circular motion instead of realistic physics-based movement. It received a score of 18.
  • CLI Tool in Rust & Blender Script: Both the Rust CLI tool and Blender script generation attempts failed.
  • Riddle Solving: The model incorrectly answered a riddle, identifying “salt” as the answer to a riddle whose answer is “smoke,” and surprisingly attempted to generate an HTML file for the response.

These benchmarks resulted in a 32nd position on the leaderboard, below Gemini 3 Pro but above GPT-5.2. The speaker acknowledges that this benchmark may not provide a complete picture of the model’s capabilities.

Agentic Capabilities: Initial Observations

Preliminary testing of Gemini 3 Flash’s agentic capabilities revealed a persistent issue common to Gemini models: improper tool calling. Specifically, the model incorrectly offered multiple-choice options even when presented with a simple greeting ("hi"). This behavior indicates the model can perform tool calling, but lacks the “sense” to use it appropriately. The speaker contrasts this with GLM 4.6 and Miniax, which handle tool calling much more effectively. The speaker explains that the multiple-choice tool is intended for proposing changes with multiple implementation options, not for responding to basic greetings. He intends to publish a full video on agentic benchmarks tomorrow.

Cost and Comparison

The speaker notes that while Gemini 3 Flash is a good model, the price point may be relatively high and plans to compare it to other similarly priced models in an upcoming video.

Conclusion

Gemini 3.0 Flash presents a promising option for applications requiring fast and cost-effective multimodal processing. While it demonstrates strong performance in certain areas like SVG generation, it struggles with more complex tasks like code generation and agentic reasoning, particularly regarding sensible tool calling. The speaker emphasizes the need for further evaluation, especially in comparison to other models within the same price range, to determine its overall value proposition. As stated by the speaker, “Overall, it’s pretty cool.”

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.