Kimmy K 2.5 Agentic Testing: A Detailed Analysis
Key Concepts:
- Mixture of Experts (MoE): A neural network architecture utilizing multiple “expert” networks, activating only a subset for each input, increasing capacity without proportional computational cost.
- Agentic Testing: Evaluating a large language model’s ability to autonomously plan and execute complex tasks, often involving tool use (like APIs and web browsing).
- Multimodal Model: A model capable of processing and understanding multiple data types, such as text and images.
- Drizzle ORM: An Object-Relational Mapper for TypeScript, simplifying database interactions.
- SQLite: A self-contained, serverless, zero-configuration, transactional SQL database engine.
- FTS5: Full-Text Search version 5, a module in SQLite for efficient text searching.
- SSE (Server-Sent Events): A server push technology enabling real-time updates from the server to the client.
- Kilo Code/Verdant: Platforms used for running and evaluating LLM agents, offering features like tool access and background task execution.
1. Introduction & Model Overview
The video details agentic testing performed on Kimmy K 2.5, the latest iteration of the Kimmy model. Initial testing showed promising results, prompting a deeper dive into its capabilities, particularly its newly added multimodal functionality. Kimmy K 2.5 retains the trillion parameter MoE architecture of K2, activating 32 billion parameters at any given time. A significant update is the inclusion of native vision capabilities, trained on approximately 15 trillion mixed visual and text tokens.
2. Performance Trade-offs: Writing vs. Coding
A notable observation is a perceived trade-off in Kimmy K 2.5’s capabilities. While it appears to have improved in coding, its writing abilities have reportedly been “lobotomized” or diminished. The speaker notes Kimmy was previously exceptional at writing, surpassing other models, but this strength seems to have been sacrificed for coding proficiency. There is speculation that the model was trained on output from Claude, potentially explaining this shift.
3. Agentic Coding Capabilities: Detailed Examples
The core of the video focuses on evaluating Kimmy K 2.5’s agentic coding skills through several practical tasks:
- Expo Movie Tracker App (TMDB API & Local Storage): Kimmy successfully created a functional app utilizing the TMDB API and local storage for reviews. The model demonstrated a methodical approach: creating a to-do list, installing dependencies (async storage, expo blur, date fns), and implementing proper TypeScript types and separation of concerns. The resulting app features inner pages, review addition, and a calendar view of watch history, comparable to Opus in functionality but at a significantly lower cost.
- Go-Based Terminal Calculator (Bubble Tea): The model generated clean and well-structured Go code using the Bubble Tea framework and Lip Gloss for styling. It implemented a proper model struct to manage calculator state, a full button grid layout with color-coding, and handled edge cases like division by zero. The code compiled successfully at a cost of only $0.05.
- Godot Game Step Calculator: Given a basic Godot game boilerplate, Kimmy successfully implemented a step calculator with a progress bar. It analyzed existing code, created a to-do list (step tracking, progress bar UI, life bar, fall damage, jump health cost, settings), and integrated the step tracking with player physics. Configurable step targets were also implemented.
- Spelt Kanban App (SQLite): Kimmy built a Kanban app with authentication, drag-and-drop functionality, and offline support, utilizing Kit with Better SQLite3 for persistence. The code structure was solid with TypeScript types and prepared statements for database queries. A minor SQLite syntax error (double quotes vs. single quotes in datetime function) was quickly identified and corrected. The authentication system employed secure HTTP-only session cookies.
- Failed Attempts: Stack Overflow Q&A, Tori App, Open Code Repo: Attempts to create a Stack Overflow-style Q&A site, the Tori app, and contribute to the Open Code Repo were unsuccessful. These failures were attributed to excessive complexity, combining NUX, Drizzle, SSE, I18N, and FTS5 search, leading to TypeScript errors and SQL syntax issues. Opus, while also facing challenges, was able to achieve partial success with finagling.
4. Performance Benchmarking & Cost Analysis
Kimmy K 2.5 achieved a score of 69.0 with Kilo Code, placing it fifth on the leaderboard. Opus 4.5 with Kilo Code leads with 77.1, followed by Opus 4.5 Ultrathink with Verdant (75.1), Gemini 3 Pro Preview with Kilo Code (71.4), and Codebuff (69.7).
Crucially, Kimmy K 2.5’s average cost per task is only $0.05, compared to Opus’s $6.86. The total cost for all seven evaluations was $3.51 for Kimmy versus $48 for Opus – a 14x cost reduction for a roughly 10% performance difference.
5. Model Behavior & Tool Usage
The speaker highlights Kimmy K 2.5’s structured approach to problem-solving, often creating a to-do list before tackling complex tasks. The model effectively utilizes browser tools for verification, building applications and then opening them in a browser to check rendering and functionality, taking screenshots, and making corrections based on visual analysis. This self-verification loop is considered a key strength of AI coding. While Kimmy K 2.5 takes longer to complete tasks than Opus, it demonstrates comparable performance in TypeScript and JavaScript.
6. Future Outlook & Conclusion
The speaker anticipates Kimmy K 2.5 becoming available on Verdant, enabling background task execution and code review integration. Overall, the model is considered a strong contender, potentially the best agentic model currently available, on par with Opus and GLM. The addition of vision capabilities further enhances its value.
Quote: “It’s a really good trade-off depending on your use case.” – referring to the cost-performance ratio of Kimmy K 2.5 versus Opus.
Synthesis: Kimmy K 2.5 represents a significant advancement in open-source agentic coding models. While a trade-off exists between writing and coding abilities, its impressive coding performance, coupled with its remarkably low cost and effective tool usage, makes it a compelling option for developers and researchers. Its structured approach and self-verification capabilities position it as a powerful tool for automating complex coding tasks.
AI summaries can miss context or contain errors. Check important details against the original video.