Sakana Fugu (Fully Tested - V/S Fable): UHM... REALLY?
By AICodeKing
Key Concepts
- Multi-Agent Orchestration: A system that routes tasks to various specialized AI models, coordinates their efforts, verifies outputs, and synthesizes a final response.
- Learned Model Router: An intelligent layer that automatically selects the most appropriate model for a specific task.
- Frontier Models: Top-tier, large-scale AI models (e.g., Fable, Mythos, Opus) that represent the current state-of-the-art in performance.
- Orchestration Tokens: Billable units consumed not just by the final output, but by the internal "work" the system performs to coordinate and verify agents.
1. Overview of Sakana Fugu
Sakana Fugu is not a traditional foundation model (like GPT or Gemini). Instead, it is a multi-agent orchestration system. It functions as a "learned model router" that takes a user prompt, decides which worker model is best suited for the task, coordinates multiple agents, verifies the results, and synthesizes the final answer.
There are two versions:
- Fugu: The faster, everyday version.
- Fugu Ultra: A performance-focused version designed for complex tasks using a deeper pool of expert agents.
The primary value proposition is to provide "Fable-level" performance without relying on a single, potentially restricted or unavailable, frontier model.
2. Benchmark Analysis
The video highlights a discrepancy between marketing claims and benchmark data. While Fugu Ultra performs competitively in some areas, it is not a "Fable killer."
- Competitive Results: Fugu Ultra performs well on Terminal Bench 2.1 (82.1 vs. Fable 5’s 89.8) and GPQA Diamond (95.5 vs. Mythos 94.6).
- Clear Deficits: Fugu Ultra falls behind in SWEET Bench Pro (73.7 vs. Fable 5’s 80) and Humanity’s Last Exam (50 vs. Fable 5’s 53.3).
- Caveats: Sakana acknowledges that Fable and Mythos are not in their agent pool (due to lack of public access) and that benchmark scores for competitors are provider-reported.
3. Practical Testing and Performance
The reviewer tested Fugu Ultra against practical coding and simulation tasks, finding the results underwhelming compared to top-tier models:
- Simulators/Games: The "elevator simulator" and "bow and arrow simulator" lacked polish, had incorrect physics, and suffered from poor user experience.
- 3JS Tasks: While the model could generate basic structures (like a folding table), it failed to implement functional mechanics (e.g., the folding/unfolding logic), which is the core requirement of the task.
- Visual/SVG Tasks: The model struggled to produce high-quality, "cute," or polished visual outputs compared to specialized models like GLM.
4. Key Arguments and Perspectives
- Marketing vs. Reality: The reviewer argues that Sakana’s marketing creates a false impression of a new, superior foundation model, whereas the product is actually an orchestration layer.
- The "Router" Problem: If the system simply routes to existing models like Opus, the user is essentially paying for an orchestration layer that may not add "extra magic" or significant value over using a strong model directly.
- Efficiency Concerns: The "verify and rewrite" process is described as potentially tedious and time-consuming. If the system is fast, it likely isn't doing deep verification; if it is doing deep verification, it becomes expensive and slow.
5. Pricing and Economic Considerations
- Hidden Costs: While the input/output pricing ($5.00/$30.00) looks competitive, the "orchestration tokens" are billable. This means complex tasks that require heavy coordination will cost significantly more than the base price suggests.
- Recommendation: The reviewer suggests that for app building, 3JS, and simulation tasks, it is more cost-effective and reliable to use a strong, direct model or a cheaper alternative like GLM rather than paying for an orchestration layer that does not consistently improve the final output.
6. Synthesis and Conclusion
Sakana Fugu represents an interesting engineering direction in "learned orchestration," which may prove useful for specific, high-level research or multi-step analysis tasks. However, as a replacement for frontier models like Fable, it currently falls short. The system functions more as an automated router than a revolutionary new model, and its real-world performance in coding and simulation tasks does not yet justify the marketing hype or the potential overhead costs.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

AI System Design: From Idea to Production - Apoorva Joshi, MongoDB
AI Engineer

When All Context Matters: Extended Cache Augmented Generation - Luis Romero-Sevilla, Orbis
AI Engineer

Bypassing the Multimodal Tax: Hybrid RAG, SQL RRF & UI Telemetry - Abed Matini, Ogilvy
AI Engineer

OpenClaw in Your Hand: Building a Physical AI Terminal - Lech Kalinowski, Callstack
AI Engineer

GPT 5.6 Mythos Level Intelligence
Prompt Engineering

GPT 5.6 SOL: TBH, IT'S OKAY.. I have SERIOUS CONCERNS.
AICodeKing

Sakana Fugu Ultra BEATS Fable 5 & GPT-5.5? (Fully Tested)
WorldofAI