Anthropic Claude Sonnet 4.6: A Detailed Analysis
Key Concepts:
- Claude Sonnet 4.6: The latest iteration of Anthropic’s Sonnet model, featuring a 1 million token context window.
- Agentic Coding: Utilizing LLMs as autonomous agents to build and modify code.
- King Bench: A benchmark used to evaluate LLMs across general knowledge and coding tasks.
- One-Shot Tasks: Tasks requiring a single prompt and response, testing general knowledge and reasoning.
- Kilo Code: A command-line interface used for agentic evaluations with LLMs.
- Verdant: A platform for managing and monitoring multiple tasks assigned to an LLM agent.
- Context Window: The amount of text an LLM can process at once (1 million tokens in Sonnet 4.6).
- Token: A unit of text used by LLMs for processing (approximately 4 characters).
I. Overview & Anthropic’s Claims
Anthropic has released Claude Sonnet 4.6, presenting a complex upgrade. The model boasts a 1 million token context window (currently in beta) while maintaining the same pricing as Sonnet 4.5: $3 per million input tokens and $15 per million output tokens. It is now the default model for free and Pro Claude users. Anthropic claims that early testing showed a 70% user preference for Sonnet 4.6 over Sonnet 4.5 when modifying code, citing improved context reading. Furthermore, they state that 59% of users preferred Sonnet 4.6 over Opus 4.5, a significant claim for a Sonnet-tier model.
II. King Bench Performance: A Regression in One-Shot Tasks
Initial testing using the King Bench benchmark revealed unexpected results. Sonnet 4.6 scored 59% overall, a 3% decrease compared to Sonnet 4.5’s 62%. This drop positioned Sonnet 4.6 at number 12 on the leaderboard, falling behind models like Kimmy K 2.5 and Gemini 3.0 zero flash skyhawk.
A breakdown by category highlights the disparity:
- General Knowledge: A substantial regression from 40% (Sonnet 4.5) to 25% (Sonnet 4.6).
- Math Questions (Q9 & Q10): Sonnet 4.5 scored 2/20 on both, while Sonnet 4.6 improved slightly to 5/20 on both.
- Riddle Question (Q11): A significant drop from a perfect 20/20 (Sonnet 4.5) to 5/20 (Sonnet 4.6).
- Coding: Performance remained largely consistent, with both models scoring 71%. Specific coding task scores for Sonnet 4.6 were: 15/20 (3D floor plan), 15/20 (SVG panda), 10/20 (Pokeball), 15/20 (chessboard), 18/20 (3D Minecraft), 16/20 (butterfly), 13/20 (Rust CLY tool), and 12/20 (Blender script).
Notably, the cost of running the full King Bench increased from approximately $43 (Sonnet 4.5) to $80 (Sonnet 4.6), nearly doubling the expense for a model with diminished overall performance.
III. Hypothesis: A Smaller, Optimized Model
The presenter theorizes that Sonnet 4.6 may be a retrained, smaller model – potentially what was initially intended as Sonnet 5 – that was scaled down. The observed regression in general knowledge, consistent coding performance, increased cost (potentially due to more verbose processing), and overall behavior support this hypothesis. The increased cost is unusual for a model upgrade, suggesting it may be compensating for reduced raw knowledge with more extensive processing.
IV. Agentic Coding Performance: A Dominant Force
Despite the shortcomings in one-shot tasks, Sonnet 4.6 demonstrates exceptional performance in agentic coding scenarios. When paired with Kilo Code, it achieved an average score of 87.9 on the King Bench agent leaderboard, securing the top position. This surpasses models like GLM 5 plus Kilo CL I (84.1) and Opus 4.6 plus Claude Code (83.6). This means Sonnet 4.6 outperforms Opus 4.6 specifically when used as an agent for coding.
V. Real-World Coding Project Evaluations
The presenter tested Sonnet 4.6 on five complex, real-world coding projects:
- Go-based Terminal Calculator (Bubble Tea): Successfully built a fully functional calculator with approximately 370 lines of code, including grid layout, navigation, error handling, and unit tests, completed in 3.5 minutes.
- React Native Expo Movie Tracker (TMDB API): Created a complete application with a custom design system, activity heatmap, API integration, local storage, and TypeScript types, compiling without errors.
- Next.js Q&A Platform: Developed a full-stack application resembling Stack Overflow, featuring server-sent events, static regeneration, internationalization, authentication, rate limiting, voting, comments, and revision history.
- Svelte Kanban Board (Offline Mutation Queue): Built a Kanban board with offline support using IndexedDB and a mutation queue, including drag-and-drop functionality and unit/end-to-end tests.
- Cross-Platform Image Cropping/Annotation App (Tauri, Rust, TypeScript): Constructed a desktop application with image processing, system tray integration, deterministic packaging, and a full canvas renderer, passing all build and check processes without errors.
These projects demonstrate Sonnet 4.6’s ability to handle complex, multi-file applications from scratch, showcasing its planning, execution, and code organization capabilities.
VI. Cost Comparison & Practical Implications
Sonnet 4.6’s cost remains at $3 per million input tokens and $15 per million output tokens, significantly lower than Opus 4.6’s $5/$25. This cost advantage, combined with superior agentic coding performance, makes Sonnet 4.6 a compelling choice for developers utilizing coding agents.
The presenter states, “For vibe coding, I prefer Sonnet 4.6 over Opus 4.6 and Codeex.” They emphasize that Sonnet 4.6 plans more carefully, reads existing code effectively, breaks down tasks logically, and maintains consistency during complex projects.
VII. Future Testing & Conclusion
The presenter plans to further evaluate Sonnet 4.6’s performance on Verdant, a platform for managing multiple concurrent tasks.
Conclusion: Anthropic’s Sonnet 4.6 presents a trade-off. While it exhibits a regression in one-shot tasks and general knowledge, it excels in agentic coding, surpassing even Opus 4.6 in performance and offering a more cost-effective solution. The model appears to be deliberately optimized for agentic workflows, potentially at the expense of broader intelligence. Its suitability depends heavily on the intended use case: a downgrade for casual users, but a significant upgrade for developers leveraging coding agents.
AI summaries can miss context or contain errors. Check important details against the original video.