Gemini 3 Pro: It was worth the wait!!!
By Prompt Engineering
Key Concepts
- Gemini 3: The latest and most intelligent Gemini model from Google, featuring advanced multimodal reasoning and agentic capabilities.
- Agentic Capabilities: The ability of an AI model to plan, take actions, and verify outcomes, acting like a junior developer or assistant.
- Google Anti-gravity: A new agentic coding IDE from Google, designed to function as a junior developer.
- Multimodal Reasoning: The ability of an AI model to process and understand information from multiple modalities, such as text and images.
- UI Generation: The capability of an AI model to create user interfaces, particularly web applications.
- LM Arena: A benchmark for evaluating large language models, where Gemini 3 ranks at the top.
- WebDev Arena: A benchmark specifically for web development capabilities, where Gemini 3 also scores highly.
- Temporal Consistency: The ability of an AI model to maintain coherence and logic across sequential frames or events in an animation.
- SVG (Scalable Vector Graphics): A vector image format that is resolution-independent and can be animated.
- Grounding with Google Search: The ability of an AI model to use Google Search to verify information and reduce hallucinations.
- Hallucination: The tendency of AI models to generate false or nonsensical information.
- Chain of Thought: The step-by-step reasoning process of an AI model, which can be observed to understand its decision-making.
- Misguided Attention: A failure mode where an AI model focuses on irrelevant aspects of a prompt or problem, leading to incorrect solutions.
Gemini 3: A Leap in AI Capabilities
Introduction and Initial Impressions
Gemini 3 has been officially released after a week of leaks, and it is described as the most intelligent Gemini model to date. The reviewer has been impressed with its multimodal reasoning capabilities, calling them "next level." Google's focus on agentic capabilities is evident, with Gemini 3 excelling in planning, taking actions, and verifying. A significant surprise accompanying Gemini 3 Pro is the release of Google Anti-gravity, a new agentic coding IDE positioned as a "junior developer."
Benchmarks and Performance
In early access, Gemini 3 has demonstrated top-tier performance on key benchmarks:
- LM Arena: Gemini 3 achieved a score of 151, placing it at the top of the leaderboard, matching Gemini 2.5 Pro.
- WebDev Arena: The model scored 1487, also ranking it as a top model in this category.
The reviewer notes that Gemini 3 is particularly strong in UI generations, a detail that will be explored further.
Multimodal Reasoning and UI Generation Examples
Gemini 3's multimodal reasoning is showcased through an application on AI Studio. Users can upload an image, and the model generates working web apps based on its content.
- Cassette Player Example: An image of a cassette and player resulted in a functional web app where users could play and eject cassettes.
- Black Hole Simulation: Uploading an image of a black hole led to a simulation demonstrating gravity warps and the behavior of objects orbiting it, highlighting its potential for educational purposes.
- YouTube Logo Generation: The model created a web app based on the YouTube logo.
Google Anti-gravity: The Junior Developer IDE
Google Anti-gravity is presented as a revolutionary approach to software development, not just a coding assistant but a "junior developer." It allows users to:
- Plan tasks.
- Assign tasks to the IDE.
- Test tasks using its own browser.
- Review and comment on the results.
This offers a distinct coding experience compared to existing solutions. A dedicated video on Anti-gravity is planned.
Gemini 3 Testing and Real-World Applications
The reviewer shares personal testing results with Gemini 3, emphasizing its strengths in web design and animation.
- Login Page with 3D Character: A prompt to create a login page with a physics-based 3D flower character that tracks the mouse cursor resulted in a "really, really neat" output, with all animation generated from a single static image.
- Voice-to-Text App Landing Page: For a voice-to-text app landing page, Gemini 3 generated a design that "does not look like a template" and achieved this in a "single shot," a significant improvement over previous models.
- "Hello World" Animation: A complex prompt to create an animation of a crowd forming "Hello World" as the camera angle changes to a bird's view was successfully executed with temporal consistency, a feat not achieved by other models.
- Animated SVG Pedestal Fan: The model generated a playable Minecraft game and an animated SVG of a pedestal fan with speed controls.
- Personal Software Development: The reviewer is using Gemini 3 to build a fully local video editor with expected features.
Gemini 3 Pro Preview Features and Capabilities
The Gemini 3 Pro Preview offers several key features:
- Default Temperature Setting: Standard parameter for controlling randomness.
- Multimodal Inputs: Accepts images and other modalities.
- Resolution Control: Influences the number of tokens processed.
- Function Calling: Not available in the early access preview.
- Grounding with Google Search: Ability to use Google Search for context and verification.
- URL as Context: Can process information from provided URLs.
- Output Length: Up to 65,000 tokens.
- Context Window: Approximately 1 million tokens.
- Thinking Level: Options for "low" and "high" thinking processes.
Detailed Testing: Rubik's Cube Solver and Animation Comparison
- Rubik's Cube Solver: Gemini 3 successfully generated code for a Rubik's cube solver that allows users to initialize cubes of different sizes, shuffle them, and then solve them. The chain of thought demonstrated a structured approach: understanding the request, planning, and then implementing. The preview feature in AI Studio allows for immediate visualization of the generated code.
- Animation Comparison (Gemini 2.5 Pro vs. Gemini 3 Pro):
- Prompt: Create an animation of a crowd walking to form "Hello World, I'm Gemini," with a bird's eye view.
- Gemini 2.5 Pro: Generated approximately 6,000 tokens, resulting in a basic surface with no animation.
- Gemini 3 Pro: Generated approximately 3,500 tokens but produced a superior animation where characters walked randomly and then formed the requested words, demonstrating impressive temporal consistency.
Web Design and UI Generation Prowess
Gemini 3 is highlighted as a leader in graphical user interfaces (GUIs), capable of generating UIs that "don't feel like created by an AI."
- Voice-to-Text System Website: The model created a "really, really nice job" with a website for a local voice-to-text system.
- Iterative Design with Web Link: By providing a web link and asking for design iterations, Gemini 3 effectively accessed and utilized the website's content to create new designs.
- Advanced Animation Prompt: Responding to a tweet showcasing a complex animation, Gemini 3, with provided image and behavior definitions, generated a first implementation with the general behavior, and a second iteration that included other elements, demonstrating strong iterative capabilities.
Failure Cases and Limitations
Despite its strengths, Gemini 3 exhibits some limitations:
- Clock Test Failure: When presented with an analog clock image, the model confused the hour and minute hands, leading to an incorrect time reading. The chain of thought revealed it correctly identified the hands' positions but misattributed their roles.
- Misguided Attention in Riddles:
- Modified Trolley Problem: Gemini 3 correctly identified that the five people on the track were already dead and chose not to pull the lever, demonstrating advanced reasoning.
- River Crossing Problem: While it identified the goat as the core goal, it struggled with the specific modified version of the riddle, attempting to solve the original unmodified problem and providing a multi-step solution when only one step was required. The reviewer notes that the chain of thought in this case describes actions rather than deep thinking.
Personal Software Development: Text-Based Video Editor
The reviewer emphasizes the development of personal software as a key strength of Gemini 3. They are building a text-based video editor, and Gemini 3 has surpassed previous agentic tools in its capabilities.
- Features Demonstrated: The editor tracks words while a person speaks, allows for adding captions by highlighting text, includes zoom in/out functionality, and enables deletion of text segments to remove corresponding video parts. It also has the ability to remove filler words.
- Comparison to Sonnet 4.5: The reviewer states that Gemini 3 is "very close or maybe surpasses" Sonnet 4.5 based on their development experience.
Reasoning and Hallucination Testing
- Misguided Attention Questions: As mentioned with the riddles, Gemini 3 shows advanced reasoning but can still suffer from misguided attention.
- Hallucination Test (Grounding with Google Search):
- Prompt: A tech journalist with a deadline asks about a surprise Google event, Gemini 3's announcement date, and rumored capabilities, with grounding enabled.
- Response: The model generated a narrative about a "shadow release" and provided details on leaked features, seemingly confusing rumors with confirmed information. This indicates that while grounding helps, it can still be influenced by pre-existing information or leaks.
Benchmarks and Final Thoughts
- Leaderboard Performance: Gemini 3's top ranking on LM Arena (1500+) and strong performance on other benchmarks like Humanities Last Exam (37.5%), GPQA (92%), MMLU Pro (81%), and MM Video MMU (87.6%) are highlighted.
- Deep Think Version: A "deep think" version of Gemini 3 also shows breakthrough performance on benchmarks like Humanities Last Exam and GPQA without tool use.
- Vibes and Future Development: The reviewer suggests that at this stage, differentiating between models is becoming difficult. The focus should shift to what can be built on top of these powerful models. Gemini 3 is recommended for users to try, with positive "vibes."
Conclusion
Gemini 3 represents a significant advancement in AI, particularly in multimodal reasoning, UI generation, and agentic capabilities. While it exhibits some limitations, such as the clock test and occasional misguided attention, its ability to generate sophisticated UIs, create complex animations, and facilitate personal software development makes it a highly impressive and useful tool. The introduction of Google Anti-gravity further signals a shift towards AI as a collaborative development partner.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Git Rerere: The Secret Merge Feature
NeuralNine

Git Crash Course - Full Tutorial For Beginners
NeuralNine

Figma Unveils Full-Stack Canvas for the AI Era
Bloomberg Technology

Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Applications, Coding AI
Stanford Online

A Genius With Amnesia - Victor Savkin, Nx
AI Engineer

Understanding Loop Engineering
GitHub

Sakana Fugu (Fully Tested - V/S Fable): UHM... REALLY?
AICodeKing