Gemini 3.1 Pro: A Detailed Performance Analysis
Key Concepts:
- Gemini 3.1 Pro: Google’s latest large language model (LLM), marketed as a significant reasoning upgrade with a 1 million token context window and 65,000 token output limit.
- KingBench: A personal benchmark used by the video creator to evaluate LLM performance, encompassing both “oneshot” and “agentic” tasks.
- Oneshot Tasks: Tasks requiring a single prompt and response, testing general knowledge and coding ability.
- Agentic Tasks: Tasks requiring the LLM to act as an agent, utilizing tools and planning to achieve a goal (e.g., building applications).
- Kilo CLI: A framework used for evaluating agentic performance, providing access to tools and a structured environment.
- Context Window: The amount of text an LLM can process at once. A larger context window allows for more complex reasoning and understanding.
- Token: The basic unit of text processed by an LLM. Cost is often calculated per token.
- ARC AGI2: A benchmark designed to measure advanced reasoning capabilities.
- SWEBench: A suite of benchmarks focused on software engineering tasks.
1. Introduction & Initial Claims
Google recently released Gemini 3.1 Pro, a 0.1 increment update – a departure from previous 0.5 version releases. Google promotes this version as a substantial improvement in reasoning capabilities, citing a score of 77.1% on ARC AGI2, a significant increase from Gemini 3 Pro’s 31.1%. However, the video creator’s testing reveals a different story.
2. Oneshot Benchmark Results (KingBench)
Testing on the KingBench oneshot benchmark showed Gemini 3.1 Pro Preview achieved a 96% score (212/220), with 100% on general questions and 95% on coding. While a good score, it regressed compared to Gemini 3 Pro, which achieved a perfect 100% (220/220). This regression occurred despite a more than doubling in cost: Gemini 3 Pro cost $0.85 to run the benchmark, while Gemini 3.1 Pro Preview cost $1.73.
A comparative cost/performance analysis places Gemini 3.1 Pro third on KingBench for oneshot tasks:
- Claude Opus 4.6: 100% - $6.39
- Gemini 3 Pro: 100% - $0.85
- Gemini 3.1 Pro Preview: 96% - $1.73
- GLM5: 79% - $0.14
3. Agentic Benchmark Results (KingBench & Kilo CLI)
The agentic performance of Gemini 3.1 Pro was significantly worse. Using Kilo CLI, the model scored 49.2% on the agent benchmark, ranking 19th out of 46 agents, with a total cost of $4.37. This represents a substantial drop from Gemini 3 Pro Preview’s score of 71.4% (rank 7).
Top performers on the agentic leaderboard include:
- Sonnet 4.6 (Kilo Code): 87.9%
- GLM5 (Kilo CLI): 84.1%
- Opus 4.6 (Claw Code): 83.6%
- Opus 4.5 (Kilo Code): 77.1%
- Minimax M2.5: 76.6%
4. Detailed Analysis of Agentic Behavior
The poor agentic performance stems from problematic planning behavior. Gemini 3.1 Pro enters excessively long planning phases, often repeating the same ideas with slightly different wording. For example, during a Go terminal calculator task, the planning phase lasted 37 seconds, filled with redundant statements like “contemplating the design,” “mapping the layout,” and “outlining the structure.”
Furthermore, the model fails to utilize the provided tools correctly. Instead of using the Kilo CLI’s question-asking tool, it embeds clarifying questions directly into its planning response, demonstrating a fundamental misunderstanding of agentic workflows.
Specific examples of problematic behavior:
- SpeltKit Conbon Task: Over 90 seconds of planning before writing any code, with the second planning phase largely repeating the first.
- Tori Image Cropper Task: 114 seconds of planning, explicitly stating it would not write code yet, followed by more clarifying questions instead of execution.
- Coding Errors: Duplicated update methods in the Go code, resulting in compiler errors. Left-in comments indicating unfinished code.
- Package Installation Errors: Attempted to install non-existent packages (e.g., "dnd-action" instead of "spelt-dnd-action," "playword" instead of "playright").
5. Pricing & Cost-Effectiveness
Gemini 3.1 Pro costs $2 per million input tokens and $12 per million output tokens – the same as Gemini 3 Pro. However, given its lower performance, the creator argues it’s not a cost-effective choice compared to alternatives. Claude Opus 4.6, while more expensive, offers superior performance. Sonnet 4.6 and GLM5 provide even more compelling value due to their lower costs and competitive performance.
6. Industry Benchmarks vs. Real-World Performance
Google claims strong results on industry benchmarks like SWEBench Verified (80.6%), Terminal Bench 2.0 (68.5%), and APEX Agents (33.5%). However, the creator emphasizes that these benchmarks don’t always reflect real-world coding performance, as demonstrated by KingBench’s results.
7. Conclusion & Recommendations
The video creator concludes that Gemini 3.1 Pro is, in most respects, worse than Gemini 3 Pro. It regressed on oneshot tasks while increasing in cost, and its agentic performance suffered a significant drop.
Recommendations:
- Free Tier Users: Gemini 3.1 Pro is a good option if accessed through free tiers like Gemini CLI or Google Anti-gravity.
- Paid API Users: Consider alternatives like Sonnet 4.6, Opus 4.6, or GLM5 for better performance and value.
- Workflow Preference: The creator prefers Verdant with Sonnet 4.6 for daily coding tasks and Opus 4.6 for more complex tasks.
Notable Quote:
“It’s actually worse than Gemini 3 Pro in almost every way that matters.” – Video Creator.
AI summaries can miss context or contain errors. Check important details against the original video.





