Gemini 3.1 Pro (Fixed with KingMode) + GLM-5: This SIMPLE TRICK makes Gemini 3.1 PRO A BEAST!

AICodeKingAbout 5 min readFeb 24, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • Gemini 3.1 Pro: Google’s large language model (LLM) with a 1 million token context window. Initially criticized for planning inefficiencies and token wastage.
  • King Mode: A system prompt designed to improve LLM discipline, focus, and consistency by reducing “fluff” and forcing direct execution.
  • Ultraink Trigger: A component of King Mode that prompts the model to assess task complexity and choose between deep reasoning and quick execution.
  • Verdant: A platform for running and orchestrating multiple AI agents in parallel, utilizing isolated git work trees.
  • GLM5: A 744 billion parameter LLM from Z AI, excelling in backend architecture and system design.
  • Tailwind CSS: A utility-first CSS framework used for rapid UI development.
  • ARC AGI2 & SWEBench: Benchmarks used to evaluate LLM performance; ARC AGI2 measures general reasoning ability, while SWEBench assesses coding proficiency.

Gemini 3.1 Pro: A Discipline Problem Solved with King Mode

This video details a re-evaluation of Google’s Gemini 3.1 Pro LLM, following initial criticism of its performance. The core argument is that Gemini 3.1 Pro’s issues aren’t a lack of intelligence, but rather a lack of discipline in its approach to tasks. This discipline problem is effectively addressed through the implementation of “King Mode,” a system prompt developed by the video creator.

Initial Concerns & Performance Metrics

The initial assessment of Gemini 3.1 Pro highlighted several shortcomings: excessive planning time, redundant thought processes, inefficient token usage, and poor tool utilization. Specifically, the model was observed to spend up to 90 seconds planning before generating any code, repeatedly stating phases like “contemplating the design,” “mapping the layout,” and “planning the implementation” – essentially repeating the same thought process.

Despite these issues, the model possesses significant underlying intelligence. It boasts a 1 million token context window, achieving a score of 77.1% on the ARC AGI2 benchmark and 80.6% on Google’s SWEBench benchmark. This demonstrates a strong capacity for reasoning and coding, but an inability to translate that capacity into efficient action.

King Mode: The Solution

King Mode is presented as a system prompt designed to instill discipline in LLMs. It doesn’t increase intelligence, but rather focuses the model by “stripping the fluff” and forcing it to prioritize direct execution. The “Ultraink” trigger within King Mode is crucial, prompting the model to assess task complexity and decide whether deep reasoning or quick execution is required.

Example: When tested on the “spelt kit conband” task, Gemini 3.1 Pro without King Mode spent 90 seconds planning across two phases, with the second phase largely replicating the first. With King Mode, the planning phase was reduced to 15 seconds, immediately followed by code generation. This resulted in significantly improved output quality due to the efficient use of the context window.

Gemini 3.1 Pro Excels at Front-End Development with King Mode

A surprising finding is Gemini 3.1 Pro’s strong performance in front-end development when used with King Mode. The model demonstrates a proficiency in generating clean, well-organized Tailwind CSS implementations, understanding responsive design principles, and leveraging its large context window to maintain an entire design system in memory.

Real-World Application: The video demonstrates a live example using Verdant. A prompt requesting a “modern portfolio website using Next.js14 with Tailwind” resulted in a polished, production-ready design with animated text, hover effects, and a visually appealing dark theme with purple accents. The model’s output showcased “design taste,” producing aesthetically pleasing code rather than merely functional code.

Backend Limitations & Synergistic Approach with GLM5

While Gemini 3.1 Pro shines in front-end tasks, it continues to struggle with complex backend architecture, database design, and API logic. The video creator acknowledges this limitation but proposes a synergistic solution: combining Gemini 3.1 Pro with GLM5.

GLM5, a 744 billion parameter model from Z AI, is described as an “absolute beast” at backend work and a leading performer on the agentic leaderboard, surpassing even Opus 4.6. It’s also freely available through Kilo Code.

Methodology: The proposed workflow involves utilizing Verdant to run two agents in parallel: one powered by GLM5 for backend tasks and another powered by Gemini 3.1 Pro for front-end tasks. Verdant manages the parallel execution and isolates the git work trees to prevent conflicts, allowing for seamless merging of the completed work.

Cost-Effectiveness: This approach is highlighted as being exceptionally cost-effective, as both GLM5 and Gemini 3.1 Pro are either free or very inexpensive to use.

Practical Tips for Implementation

The video provides three key tips for maximizing the effectiveness of this approach:

  1. Inject King Mode into Verdant Project Rules: This ensures that the prompt applies to all agents, benefiting both GLM5 and Gemini 3.1 Pro.
  2. Use the Ultraink Prefix for Complex Tasks: This triggers the model’s focused execution mode for more demanding prompts.
  3. Provide Descriptive Front-End Prompts: Detailed visual direction yields better results from Gemini 3.1 Pro.

Conclusion

The video demonstrates that Gemini 3.1 Pro, while initially disappointing, can be significantly improved through the application of King Mode. The model’s strengths in front-end development, combined with the backend capabilities of GLM5 and the orchestration provided by Verdant, create a powerful and cost-effective AI development workflow. The key takeaway is that addressing the discipline of LLMs is often more impactful than simply increasing their size or complexity. The link to the King Mode prompt is provided in the video description for easy implementation.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.