Claude Opus 4.5 (Fully Tested): Anthropic REALLY COOKED with this model! #1 on my Agentic Tests!
By AICodeKing
Key Concepts
- Claude Opus 4.5: Anthropic's new flagship model, optimized for coding, agents, and real-world computer use.
- Cost-effectiveness: Significant price reduction for Opus 4.5 compared to previous versions, making it more accessible for daily tasks.
- Benchmarking: Performance evaluation across various coding and reasoning tasks, comparing Opus 4.5 against Sonnet 4.5 and competitor models.
- Agentic Benchmarks: Evaluation of the model's ability to perform complex, multi-step tasks using agents.
- Frontend vs. Backend: Distinction in model strengths, with Opus 4.5 excelling in backend and debugging, while Gemini 3 is stronger in frontend development.
- Context Window and System Prompt: Factors influencing model performance, with limitations noted in Claude Code and improvements observed when using external platforms like Kilo.
Claude Opus 4.5 Launch and Pricing
Anthropic has released Claude Opus 4.5, their latest flagship model, specifically designed for coding, agentic capabilities, and practical computer usage. This release is positioned as Anthropic's response to models like Gemini 3 Pro. A significant aspect of this launch is the substantial price reduction for the Opus model. The new pricing is $5 per million tokens for input and $25 per million tokens for output, a considerable decrease from the previous $15 (input) and $75 (output) per million tokens. This makes Opus 4.5 much more cost-friendly and potentially usable for everyday tasks, provided sufficient budget. Anthropic's emphasis on strategies to reduce costs, such as optimizing context length, further indicates a focus on real-world applicability.
Performance Benchmarks
Claude Opus 4.5 demonstrates significant improvements across various benchmarks:
- Coding Problems (Aider Polyglot): Opus 4.5 achieves approximately 89.4%, a notable jump from Sonnet 4.5 at 78.8%.
- Semantic Coding (SBench Verified): Opus 4.5 scores 80.9%, surpassing Sonnet 4.5 (77.2%) and Opus 4.1 (74.5%).
- Terminal Bench 2.0: Performance increases from 46.5% for Opus 4.1 to 59.3% for Opus 4.5.
- Multilingual Coding (SBench Multilingual): Opus 4.5 leads Sonnet 4.5 and Opus 4.1 across multiple languages (C, Go, Java, JS/TS, PHP, Ruby, Rust) with higher pass rates and consistent error bars.
- Long-Term Coherence (Vending Bench): The score improves from $3,849.74 for Sonnet 4.5 to $4,967.6 for Opus 4.5, indicating better sustained performance over extended tasks.
- Agents (Browse Comp Plus): Opus 4.5 reaches 72.9% when paired with tool result clearing, memory, and context resetting, compared to Sonnet 4.5 at 67.2%.
- Safety (Concerning Behavior Metric): Opus 4.5 shows a reduction to around 10%, lower than Sonnet 4.5 and competitor frontier models.
- Prompt Injection Susceptibility: Opus 4.5 exhibits the lowest susceptibility at K=1 queries (4.7%) compared to Sonnet 4.5 (7.3%) and other models. At K=100, it scores 63.0%, again outperforming Sonnet 4.5 (72.4%).
- Reasoning Heavy Tasks (Non-Coding):
- AI2 Verified: Opus 4.5 achieves 37.6%, a significant improvement over Sonnet 4.5's 13.6%.
- GPQA Diamond: Scores 87.0%.
- Visual Reasoning (MMU Val): Achieves 80.7%.
Sponsor: Augment Code
The video briefly highlights Augment Code, an enterprise-grade AI assistant for engineering teams working with large codebases. Key features include:
- Proprietary Context Engine: Delivers millisecond-relevant snippets across massive monorepos.
- Real-time Repo Feeding: Processes entire repositories, even millions of lines, into the best available model.
- Seamless Integration: Works with VS Code, Jet Brains, Vim, and Cursor without editor switching.
- Security: Secure by default, no code training, and supports customer-managed encryption keys.
- Pricing: Pay-per-message, no seat licenses or complex token math.
- New Features: Remote agents for launching, monitoring, and merging pull requests from cloud workers. A 14-day free trial is available at augmentcode.com.
Personal Benchmark Testing
The presenter also tested Claude Opus 4.5 on their own benchmarks, with mixed results for non-agentic tasks:
- Floor Plan Generation: "Fine," but not the best.
- SVG Panda Holding Burger: "Pretty bad."
- Pokeball in 3JS: "Quite good," functional but desired a better background.
- Chessboard with Autoplay: "Doesn't work," rated as "not good."
- Minecraft Game Clone (Kandinsky Style): "Really good," one of the best generations.
- Majestic Butterfly Flying Simulation: "One of the best generations," realistic physics and one-shot generation.
- CI Tool in Rust: "Pretty awesome."
- Blender Script: "Really good," including lighting and camera setup.
- General Questions (Math/Riddle): Scored 74%, which is below Gemini 3's worst checkpoint and significantly below Gemini 3 Pro's official checkpoint.
Agentic Benchmark Performance
The presenter's agentic benchmarks show Opus 4.5 performing exceptionally well, particularly when using platforms like Kilo Code:
- Expo Mobile Tracker App: "Nails it," one of the best generations, with a good UI and functionality.
- Go Terminal Calculator with Bubble T: "Nails this prompt," "amazingly good," and works well.
- Godo Game: "Works kind of well," functional but with minor UI placement issues (health bar, step calculator).
- Open Code Task (Add SVG Command): "Nailed this," completed in one go.
- Spelt App: "One of the best generations," with full functionality (login, signup, boards, tasks) and SQLite database integration.
- Next App: "Worked really well."
- Tari App: "Worked well."
These agentic tasks led to Opus 4.5 scoring the number one position on the Agentic leaderboard.
Cost Comparison and Model Strengths
Despite its strong performance, Opus 4.5 remains more expensive than Gemini 3. Gemini 3 scores 71.4% for $8, while Opus 4.5 scores 77.1% for $48. This represents a significant price jump.
- Opus 4.5 Strengths: Backend development, debugging, and complex agentic tasks.
- Gemini 3 Strengths: Frontend development.
The presenter suggests a hybrid approach: using Gemini 3 for frontend tasks and Opus 4.5 for backend and more complex operations, potentially building a functional draft with Opus and then refining the frontend with Gemini.
Limitations and Conclusion
While Anthropic is seen as "back" with this model, the price is still considered high. The presenter notes that Claude Code itself can limit model capabilities due to restricted context windows and suboptimal system prompts. However, using Opus 4.5 with platforms like Kilo Code significantly improves its performance.
Overall Takeaway: Claude Opus 4.5 is a powerful leap forward, especially for coding and agentic tasks, with a more accessible price point. However, its cost remains a factor, and its frontend capabilities are weaker compared to competitors like Gemini 3. The model's true potential is unlocked when used with external tools that optimize its context and prompt handling.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

GPT 5.6 SOL: TBH, IT'S OKAY.. I have SERIOUS CONCERNS.
AICodeKing

GLM 5.2 Is INSANE. Better than Claude Fable 5?
Zubair Trabzada | AI Workshop

GLM-5.2 (Fully Tested): I got EARLY ACCESS & This MODEL is CRAZY!
AICodeKing

20 days of compute vs 7 hours: rethinking what state-of-the-art means — Bertrand Charpentier, Pruna
AI Engineer

API vs Subscriptions vs Local: I Measured Intelligence Per Dollar.
Eduards Ruzga

Gemini 3.5 Flash In Arena! POWERFUL, Cheap, & Fast NEW AI Model! (Fully Tested)
WorldofAI

Qwen 3.6 Max: NEW Powerful AI Model EVER! Beats Opus 4.5, Gemini 3, Deepseek v4! (Fully Tested)
WorldofAI