Is Gemini 3 Really the Best AI Ever?
By Cole Medin
Key Concepts
- Gemini 3: Google's latest large language model (LLM).
- Benchmarks: Standardized tests used to evaluate LLM performance.
- AI Coding: The use of AI models to assist in software development.
- Developer Productivity: The efficiency and output of software developers.
- Anti-gravity: Google's new AI IDE (Integrated Development Environment) that integrates with Gemini 3.
- Claude Sonnet 4.5: A previously considered top LLM for coding.
- Kleinbench: A new framework for evaluating LLMs based on real-world engineering tasks.
- Agentic Coding: AI coding assistants that can perform tasks autonomously.
- System Prompt: Instructions embedded within an AI system to guide its behavior.
- Open Source Repositories: Publicly accessible code repositories, like those on GitHub.
Gemini 3: Benchmarks vs. Real-World Application
The release of Google's Gemini 3 has generated significant hype, with impressive benchmark results suggesting a substantial leap in LLM capabilities. However, the video argues that these benchmarks, while showing "insane" jumps, may function more as marketing material. A key concern is that LLMs are increasingly trained to excel at these specific test formats, potentially masking their true performance in practical, real-world scenarios.
This disconnect is highlighted by a study showing a flat line in developer productivity despite the apparent explosion in LLM capabilities. The presenter's personal experience with AI coding assistance supports this, suggesting that improvements might be more attributable to the tools and systems built around LLMs rather than the LLMs themselves. For instance, the presenter notes that even older models like Claude Sonnet 3.5, when integrated into their current AI coding system, yielded significant development speed improvements compared to using no AI assistance. This raises the question of how much of the perceived LLM power is due to the underlying model versus the surrounding infrastructure.
Anti-gravity: A Case Study in Integrated AI Development
Google's new AI IDE, Anti-gravity, serves as a practical example of how integrated tools can enhance LLM performance, particularly for frontend development. While Gemini 3, by default in Anti-gravity, appears to outperform models like Claude Sonnet 4.5 in design and frontend building, the video questions whether this is solely due to Gemini 3's inherent capabilities or the specific tools within Anti-gravity.
A notable feature of Anti-gravity is its Google Chrome integration, allowing the AI coding assistant to autonomously navigate websites, validate them visually, and offer suggestions for improvement. This leverages Gemini 3's vision capabilities. The presenter demonstrates this by having the AI analyze their codebase, start up a frontend, scroll through it, and provide design recommendations. The AI can even record its actions, allowing users to watch its autonomous navigation and interactions with the website.
While the presenter encountered an "overload error" requiring repeated prompts, the overall experience with this integrated capability was highly impressive. The autonomous website navigation, while not entirely new (comparable to tools like Playwright or Stage Hand MCP servers), is optimized within Anti-gravity. Features like waiting for page rendering before capturing screenshots are baked into the system prompt for this integration, showcasing a sophisticated design. Anti-gravity also offers an "agent manager mode" that abstracts away the code, focusing on conversational interactions with coding assistants and enabling parallel task execution across repositories.
The Problem with Synthetic Benchmarks and the Rise of Kleinbench
The core problem identified is the discrepancy between LLM performance on synthetic benchmarks and their actual utility in real-world applications, especially in AI coding. Traditional benchmarks often focus on isolated tasks, like solving LeetCode-style puzzles (e.g., reversing a linked list), which do not reflect the complex, iterative nature of actual software development.
As a solution, the video introduces Kleinbench, a framework that aims to evaluate LLMs based on real engineering tasks. Kleinbench leverages publicly available open-source repositories (like those on GitHub) to observe how code changes over time as real engineers work on actual tasks with AI coding assistants.
Key purposes of Kleinbench:
- Reliable Evaluation: Replaces synthetic benchmarks with real engineering tasks. Users opt-in, and their prompts and workflows on open-source repositories are tracked.
- Documentation of System Gaps: By standardizing and publishing controlled environments for evaluation, Kleinbench helps identify areas where AI coding assistance can be improved.
- Training Data for Research: The collected data will serve as training data for future research and fine-tuning of LLMs.
Kleinbench's Methodology:
The framework requires three pieces of information for each real engineering task:
- The starting snapshot of the repository.
- The prompt given to the AI.
- The end state, which is the code that was actually committed to the open-source repository.
While seemingly simple, Kleinbench faces challenges in accounting for the diverse tools and systems engineers use. However, by standardizing these "agentic engineering environments," it aims to provide valuable real-world data for understanding which LLMs are best suited for specific projects. The presenter emphasizes that the specific tool (Kleinbench) is less important than the shift towards evaluating LLMs on real tasks.
Conclusion and Future Directions
The video concludes by reiterating that while Gemini 3 is a powerful LLM, its true capabilities are difficult to ascertain solely from benchmarks due to the influence of surrounding tools and systems. The introduction of frameworks like Kleinbench signifies a crucial shift towards evaluating LLMs on practical, real-world engineering tasks. This approach is seen as the future direction for understanding and selecting the most effective LLMs and tools for creation.
The presenter also announces a live stream on November 29th at 9:00 a.m. Central Time, where they will be giving away their remote agentic coding system.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Frontier results, on device - RL Nabors, Arize
AI Engineer

We Cut 94% of AI Coding Tokens With a Local Code Index - Rajkumar Sakthivel, Tesco
AI Engineer

FULLY FREE GLM-5.2 + Z-Code: This is ACTUALLY GOOD!
AICodeKing

FULLY FREE Unlimited API + OpenCode: MiniMax M3,Step 3.7 Flash,Nemotron 3 Ultra,GLM,Kimi!
AICodeKing

GLM-5.2 (Fully Tested): I got EARLY ACCESS & This MODEL is CRAZY!
AICodeKing

WTF Is an "AI Agent Loop"? Genius or Hype?
Greg Isenberg

Claude Mythos 5 LEAKED & IS Coming Sooner Than Expected & GPT-5.6 Checkpoint Out! Huge AI News!
WorldofAI