Key Concepts
Claude 4 Opus, Claude 4 Sonnet, coding models, reasoning, agentic workflows, Sway Bench, Terminal Bench, long-running tasks, multifile code understanding, debugging, memory, hybrid mode thinking, tool use, parallel tool execution, improved memory, Claude Code, API capabilities, code execution tool, MCP connector, files API, prompt caching, software engineering, benchmark scores, pricing, long-term reasoning, advanced memory, thinking summaries, dev mode, autonomous agents, multi-step tasks, end-to-end app generation, long context workflows, prompt engineering, chatbot, API access, open router.
Claude 4 Opus and Sonnet: New Coding Models from Anthropic
Introduction
Anthropic has released Claude 4 Opus and Claude 4 Sonnet, the next generation of their AI models, designed to set new standards in coding, reasoning, and agentic workflows.
Claude 4 Opus: The Flagship Model
- Performance: Claude 4 Opus is positioned as the world's best coding model, achieving a score of 72.5% on the Sway Bench and 43.2% on the Terminal Bench.
- Capabilities: It excels in complex reasoning, coding, and agentic workflows. Opus is designed for long-running tasks, maintaining focus for hours. It supports deep multifile code understanding, editing, and debugging, and integrates with tools like Cursor, Replit, and Block.
- Memory: Opus maintains memory across tasks, storing key information and local files for long-term coherence. An example provided is creating a new navigation guide while playing Pokémon Red, where Claude autonomously logs game-critical notes.
- Thinking Summaries: Claude 4 Opus uses thinking summaries to keep its thoughts readable, condensing them only when necessary.
- Use Cases: Ideal for developers, researchers, and power users needing high performance and full control. Best for end-to-end app generation, long context workflows, and prompt engineering.
- Key Features:
- Reliable Long-Term Reasoning: Improves task fidelity by 65% over Sonnet 3.7 by avoiding shortcut-taking.
- Advanced Memory Through Local Files: Creates and updates memory files, retaining critical information across long, multi-step workflows.
- Thinking Summaries + Dev Mode: Condenses long thought chains for readability, while developer mode provides raw chain of thought for debugging and prompt engineering.
Claude 4 Sonnet: Balancing Performance and Speed
- Performance: Claude 4 Sonnet is a major upgrade from Sonnet 3.7, scoring 72.7% on the Sway Bench test.
- Capabilities: It strikes a balance between performance and speed.
- Hybrid Mode Thinking: Offers a hybrid mode where users can switch between instant replies and extended thinking for deep reasoning.
- Shared Improvements: Shares key improvements with Opus, such as reduced shortcut behaviors, tool use, and reasoning tool switching.
- Use Cases: Offers solid performance at a lower latency and cost.
New Features Introduced
- Hybrid Mode Thinking: Allows switching between instant replies and extended thinking for deeper reasoning.
- Tool Use: Alternates between reasoning and external tools like web search.
- Parallel Tool Execution: Enables parallel execution of tools.
- Improved Memory: Enhanced memory capabilities from file storages.
- Claude Code: A tool released earlier in the year, now generally available (GA) with native VS Code and JetBrains extensions. It can run as a background task with GitHub actions and be used as an SDK for custom agents.
- New API Capabilities: Includes code execution tool, MCP connector, files API, and prompt caching (up to 1 hour).
Benchmark Results
- Sway Bench Verified Test: Claude Opus 4 and Sonnet 4 lead with accuracies of 72.5% and 72.7%, respectively, outperforming models like OpenAI's CodeX 1, O3, and GPT-4.1.
- Parallel Test Time Compute: With parallel test time compute enabled, Opus 4 reaches 79.4%, while Sonnet 4 tops at 80.2%.
- Agentic Coding and Tool Use: The models outperform others in agentic coding, terminal coding, and tool use benchmarks.
Pricing
- Claude Opus 4: $15 per 1 million input tokens and $75 per 1 million output tokens (expensive).
- Claude Sonnet 4: Same pricing as Claude 3.7: $3 per 1 million input tokens and $15 per 1 million output tokens.
Real-World Examples and Testing
- Browser Agent: Claude 4 Opus one-shotted a full browser agent with API and front-end access in a single prompt.
- Personal Finance Tracking App: Both Opus and Sonnet were tested to build a responsive web page for tracking monthly income and expenses. Opus generated a more comprehensive app with features like night mode and data visualization.
- TV Channel Simulator: Sonnet was tested on coding a TV channel simulator, generating creative and visually diverse channels. Opus also generated a TV simulator, with higher quality animations.
- SVG Butterfly: Opus created a more accurate and aesthetically pleasing SVG representation of a butterfly compared to Sonnet.
- Tetris Game: Both models were tasked with creating a Tetris game. Opus generated a version with animations and a scorecard, while Sonnet created a functional Tetris game.
Accessing the Models
- Chatbot (Claude AAI): Access both models through their chatbot.
- Console: Test the models within their console.
- API: Access the models via their API.
- Open Router: Access the models through Open Router.
- Cursor: Integrate the models into Cursor using the Anthropic API provider.
Conclusion
Claude 4 Opus and Sonnet represent significant advancements in coding and AI capabilities. Opus excels in complex tasks requiring long-term reasoning and memory, while Sonnet offers a balance of performance and speed. The models are accessible through various platforms, making them valuable tools for developers, researchers, and power users. The presenter was impressed with the coding capabilities of both models and is happy to see Anthropic back with a new model.
AI summaries can miss context or contain errors. Check important details against the original video.





