Gemini 3.0 Flash (Agentic Tests w/ Antigravity & CLI): In-depth agentic test of Gemini 3 Flash.

AICodeKingAbout 5 min readDec 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Gemini 3 Flash: Google’s recently launched large language model (LLM).
  • Agentic Benchmarks: Evaluating LLMs on complex tasks like app building and editing, requiring multiple steps and tool use.
  • Anti-Gravity: A free platform for interacting with Gemini models, demonstrating superior performance in the tests.
  • Gemini CLI: Google’s command-line interface for Gemini, currently requiring a waitlist and showing significantly lower performance.
  • LSP (Language Server Protocol): A protocol allowing code editors and IDEs to communicate with language servers, potentially aiding code generation.
  • Hallucinations: Instances where an LLM generates factually incorrect or nonsensical information.
  • Tool Calling: The ability of an LLM to utilize external tools or APIs to accomplish tasks.
  • ShadCN UI: A collection of re-usable components built using Radix UI and Tailwind CSS, potentially enforced by Anti-Gravity’s system prompt.

Gemini 3 Flash: Agentic Benchmarks & Platform Performance Analysis

This video details a performance evaluation of Google’s Gemini 3 Flash model, specifically focusing on its capabilities in agentic benchmarks – tasks involving app building and editing. The initial evaluation placed Gemini 3 Flash at 20th position overall with a 47% success rate, a 7% improvement over previous evaluations. However, the core of the video centers on comparing its performance across different implementation platforms: Anti-Gravity and Gemini CLI.

Platform Comparison: Anti-Gravity vs. Gemini CLI

The most striking finding is the significant disparity in performance between Anti-Gravity and Gemini CLI. While both platforms offer free access to Gemini 3 Flash, their results diverge dramatically.

  • Anti-Gravity: Consistently delivers impressive results, often exceeding expectations. The presenter describes its performance as “sorcery,” particularly in complex tasks. It demonstrates strong context retention, reduced hallucinations, and effective tool calling. The agentic contraption within Anti-Gravity is able to run for extended periods (up to an hour in one test), allowing for more thorough task completion. It appears Anti-Gravity may enforce a system prompt utilizing ShadCN UI, contributing to the polished output.
  • Gemini CLI: Exhibits significantly poorer performance, described as “bad” and “halfbaked.” The agentic contraption struggles to maintain context and often terminates prematurely. The quality of generated code and UI is noticeably lower, with the presenter noting a decline in performance over time.

Application-Specific Results

The evaluation involved several application-building tasks, revealing further nuances in Gemini 3 Flash’s performance:

  • Go TUI Calculator: Anti-Gravity “nails it,” producing well-structured and functional code. Gemini CLI generates a poorly designed and incomplete application.
  • Conban App (Spelta): Gemini CLI produces a visually outdated application (resembling designs from the 2000s), despite functional correctness. Anti-Gravity generates a significantly superior, polished application, potentially surpassing the performance of the Opus model.
  • Tar App: Both platforms failed to generate a functional application, struggling with the combination of Next.js and Rust code. The model became confused about the appropriate syntax.
  • Movie Tracker App: Results were mixed. Anti-Gravity produced a partially functional app with issues in the homepage and header. Gemini CLI failed to generate a working application.
  • Open Code Task: Both platforms performed poorly, resulting in a ninth-place ranking for Anti-Gravity and a 21st-place ranking for Gemini CLI.

Underlying Factors & Observations

The presenter identifies several potential factors contributing to the performance differences:

  • Agentic Contraption Quality: Gemini 3 Flash’s performance is heavily reliant on the quality of the surrounding agentic contraption. A weak or limited contraption (as seen in Gemini CLI) hinders the model’s capabilities.
  • Context Retention & Hallucinations: Gemini 3 Flash demonstrates improved context retention and reduced hallucinations compared to previous models.
  • LSP Support: The presenter speculates that Language Server Protocol (LSP) support may play a role in code generation quality, but questions why this wouldn’t translate to improvements in the Gemini CLI frontend.
  • System Prompts: Anti-Gravity appears to enforce a system prompt that encourages the use of ShadCN UI, leading to more visually appealing and consistent designs.

Limitations & Generalization

The presenter acknowledges that Gemini 3 Flash has limitations:

  • Language & Framework Support: The model struggles with niche languages and frameworks, often incorrectly generalizing Next.js syntax as general React syntax, leading to unfixable errors.
  • Project Consistency: While capable of producing excellent results in specific instances, maintaining consistent performance across multiple projects is challenging.

Google Product Strategy & Future Plans

The presenter observes a shifting focus within Google’s AI product ecosystem:

  • Shifting Priorities: Initial focus on Project IDX and Firebase Studio has shifted to Gemini CLI and now Anti-Gravity, with the latter receiving significant attention and free tier access.
  • Neglected Products: Gemini CLI and Firebase Studio are receiving fewer updates and appear to be losing priority. Firebase Studio, in particular, is described as “scrap” with no significant updates in three months.
  • Future Testing: The presenter plans to test Gemini 3 Pro on frontend tasks and report back on the results.

Conclusion

Gemini 3 Flash demonstrates promising capabilities, particularly in frontend development, when utilized within a robust agentic environment like Anti-Gravity. Its performance is significantly hampered by the limitations of the Gemini CLI platform. The model excels at basic coding tasks but struggles with more complex or niche technologies. Despite its limitations, Gemini 3 Flash offers substantial value, especially considering its accessibility through free tiers on platforms like Anti-Gravity. The presenter emphasizes that while the model can produce impressive results, consistent performance across projects remains a challenge.

“It’s crazy to me how bad Gemini CLI is getting lately. The Agentic Contraption used to be good, but it is falling down in quality very quickly.” - Presenter on the declining performance of Gemini CLI.

“This generation is like better than Opus in my opinion… It’s insane and it’s free. I mean, what is this?” - Presenter describing the exceptional performance of Gemini 3 Flash on the Conban app within Anti-Gravity.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.