Claude Sonnet 4.5: A Deep Dive
Key Concepts:
- Claude Sonnet 4.5: Anthropic's new frontier model, emphasizing improved coding, computer use, reasoning, and STEM performance.
- Agentic Coding: The ability of an AI model to autonomously write and execute code to achieve a specific goal.
- Terminal Use: The model's proficiency in using command-line interfaces and executing terminal commands.
- Computer Use: The model's ability to interact with and utilize computer systems and tools effectively.
- Reasoning: The model's capacity to perform complex, multi-step logical inferences.
- ASL3 Protections: Safety measures implemented to prevent misuse and ensure responsible AI behavior.
- Checkpoints: A feature in Claude Code that allows users to save the state of a task and revert to it if needed.
- Claude Code: An environment designed for coding tasks using Claude models.
- VS Code Extension: A tool that integrates Claude's capabilities directly into the Visual Studio Code editor.
- Claude Agent SDK: A software development kit for building custom agentic workflows using Claude models.
- Token Pricing: The cost associated with using AI models, based on the number of input and output tokens.
1. Introduction to Claude Sonnet 4.5
Anthropic has launched Claude Sonnet 4.5, positioning it as their new frontier model and claiming it to be the best coding model currently available. The key improvements focus on enhanced computer use, longer multi-step reasoning capabilities, and superior math and STEM performance. The pricing remains consistent with Sonnet 4, at $3 for input and $15 for output per million tokens.
2. Performance Benchmarks and Comparisons
Sonnet 4.5 demonstrates significant performance gains across various benchmarks compared to previous models like Opus 4.1, Sonnet 4, GPT-5, and Gemini 2.5 Pro.
- SWE Verified (Agentic Coding): Sonnet 4.5 achieves 77.2%, surpassing Opus 4.1 (74.5%), Sonnet 4 (72.7%), GPT-5 (72.8%), and Gemini 2.5 Pro (67.2%).
- Terminal Bench (Terminal Style Coding): Sonnet 4.5 scores 50.0%, outperforming Opus 4.1 (46.5%), GPT-5 (43.8%), Sonnet 4 (36.4%), and Gemini 2.5 Pro (25.3%).
- OSWorld (Computer Use): Sonnet 4.5 reaches 61.4%, a notable increase from Sonnet 4 (42.2%) and Opus 4.1 (44.4%).
- AIM 2025 (Reasoning - Python): Sonnet 4.5 achieves a perfect 100%, exceeding Opus 4.1 (78.0), Sonnet 4 (70.5), GPT-5 (99.6), and Gemini 2.5 Pro (94.6).
- GPQA Diamond: Sonnet 4.5 scores 83.4, close to GPT-5 (85.7) and Gemini 2.5 Pro (86.4), and higher than Opus 4.1 (81.0) and Sonnet 4 (76.1).
- Multilingual MMLU: Sonnet 4.5 achieves 89.1, comparable to Opus 4.1 (89.5) and GPT-5 (89.4).
- Visual Reasoning (MM Validation): Sonnet 4.5 scores 77.8, lower than GPT-5 (84.2) and Gemini 2.5 Pro (82.0), but an improvement over Sonnet 4 (74.4).
- Finance Agent: Sonnet 4.5 achieves 55.3, surpassing Opus 4.1 (50.9), GPT-5 (46.9), Sonnet 4 (44.5), and Gemini 2.5 Pro (29.4).
- Domain Evaluations (Finance & STEM): Sonnet 4.5 with extended thinking (16k) leads in both finance (72% win rate) and STEM (69% win rate).
3. Safety and Alignment
Anthropic emphasizes that Sonnet 4.5 is their most aligned Frontier model to date, released under ASL3 protections. This is supported by a chart showing misaligned behavior scores, where Sonnet 4.5 has the lowest score among the listed models. While improved, ASL3 safeguards can still interrupt normal content in edge domains, potentially leading to false positives. In such cases, users can switch to Sonnet 4 mid-thread.
4. Claude Code and VS Code Extension
- Claude Code: The standout feature is checkpoints, allowing users to save state mid-task and roll back instantly if something breaks.
- VS Code Extension: This extension integrates Claude's capabilities directly into the Visual Studio Code editor, enabling functionalities similar to Klein within the coder. Users can install the extension, connect their Anthropic account, and utilize Claude's features within their coding environment.
5. Claude Agent SDK
The Claude Agent SDK provides the same foundation Anthropic uses for Claude Code, enabling users to build their own agent systems. This includes creating controllers and sub-agents, such as a testing sub-agent for running commands in a sandbox, a documentation sub-agent for writing summaries and updates, and a deployment sub-agent that only acts with explicit approval. The SDK allows for parallelizing tool execution, such as running multiple bash commands in CI-like flows to maximize actions per context window.
6. Strengths and Caveats
Strengths:
- Faster and more capable at real computer use and long horizon tasks.
- Checkpoints in Claude Code are a lifesaver.
- The VS Code extension keeps everything inside the editor.
- Memory and context editing reduce manual state management.
- The Agent SDK opens the door for custom agentic workflows.
- Pricing remains flat at $3 per 15 million tokens.
Caveats:
- ASL3 safeguards can still interrupt normal content in edge domains.
- Complex browser flows across Ows or weird dynamic pages may still need babysitting.
- Visual reasoning is strong, but not the highest in the field compared to GPT-5 on some metrics.
- For truly massive code bases, repo indexing and project structure still matter.
7. Conclusion
Sonnet 4.5 appears to be a meaningful upgrade, with benchmarks supporting its improved performance. The model's reliability and everyday usage are key considerations. While the presenter plans to conduct further testing, the initial assessment is positive.
AI summaries can miss context or contain errors. Check important details against the original video.





