ByteDance Seed Code (Fully Tested): ANTHROPIC is OFFICIALLY SCARED of this MODEL!
By AICodeKing
Key Concepts
- Dubau seed code: A new AI model developed by ByteDance, a Chinese company.
- Trey: An AI code editor developed by ByteDance.
- Anthropic: A company that develops AI models like Claude.
- SWEBench verified: A benchmark used to evaluate AI models' coding capabilities.
- ZenMox: A platform that provides access to various AI models, including Dubau seed code.
- Volcano API: ByteDance's API platform for its AI models.
- Agentic benchmarks: Benchmarks designed to test AI models in more complex, task-oriented scenarios.
- Token: A unit of text used by AI models for processing and generation.
Dubau Seed Code: A New Contender in AI Coding
The video introduces a new AI model from ByteDance, the company behind TikTok, named "Dubau seed code." This model is reportedly outperforming established models like Claude and GPT-5 in coding tasks, and at a significantly lower cost. This development has potentially led Anthropic to revoke Trey's access to their Claude models, a move described as "scummy."
Performance on SWEBench Verified
The presenter highlights the SWEBench verified benchmark as a key indicator of AI model performance, noting that Anthropic is particularly invested in its results. Dubau seed code, when integrated with Trey, is currently topping this benchmark, showing an approximately 8% improvement over Claude's implementation. This success is suggested as the primary reason for Anthropic's perceived insecurity and potential delay in releasing new models.
Accessibility and Cost
While Dubau seed code was recently launched publicly, it is primarily a China-focused model. Accessing it directly through ByteDance's Volcano API platform requires a Chinese mobile number. However, it can be accessed via ZenMox, a platform similar to OpenRouter, which offers a variety of models and provides free credits for testing. ZenMox also offers an Anthropic-compatible API, which the presenter used for testing.
The model is described as "insanely cheap," costing $17 and $12 per million tokens, making it approximately 15 times cheaper than other comparable models. It also supports image and video inputs, adding to its versatility.
Performance on Non-Agentic Benchmarks
The presenter conducted tests on their own non-agentic benchmarks, with mixed results:
- Floor plan: Code was correct, but the output was not visually impressive, as expected for a focused code model.
- SVG Panda with burger: The panda was recognizable, and the burger was well-rendered, though the interaction between elements was not ideal.
- Pokeball in 3JS: The visual representation was accurate in terms of colors and characteristics, but the interactive button was missing.
- Autoplay chessboard: This feature did not work, which was a disappointment.
- Minecraft and Kandinsky style: This was a standout performance, exceeding Sonnet's capabilities. The model generated mist for depth and a randomly generated map that worked, despite some limitations in movement physics. The presenter found the prompt working on such a cheap model "insane."
- Butterfly flying in the garden: The butterfly animation was smooth, and the environment was well-rendered, though the butterfly itself was not perfectly detailed.
- CLI tool in Rust: This worked fine.
- Blender script: This did not work.
Overall, on these benchmarks, Dubau seed code achieved 15th position, which is considered good, especially given its low cost and comparison to models like Gemini 3 checkpoints. The model's speed is also noted as impressive, around 80 tokens per second.
Performance on Agentic Benchmarks
Testing on agentic benchmarks, using Claude code for optimization, yielded the following:
- Movie tracker app: Did not work and was buggy.
- God game: Did not work and produced errors.
- Go Tui calculator: This was a highly successful generation, with the model "oneshotting" the task. The UI was good, and everything functioned correctly.
- Spelt app: Did not work.
- Nux app: Did not work.
- Open Code repo question: Did not work.
These results placed the model at 12th position on the agentic leaderboard, outperforming Cursor Composer but falling behind models like Kimmy and Quen Code. The presenter speculates that the use of Claude code might have negatively impacted performance, noting that the model sometimes resorted to terminal commands instead of edit diff tools. It is suggested that the model might be specifically trained for use with Trey, as it lists Trey tools when prompted. While the raw code generated is good, the tool-calling aspect appears to be a weakness.
Conclusion and Future Outlook
The presenter expresses a strong preference for cost-effective AI models that offer substantial performance, even if not at the absolute cutting edge. They believe that Chinese companies, including ByteDance with Dubau seed code, are aligning with this philosophy, similar to Minimax and GLM. The model's low cost, combined with its strong performance on coding benchmarks and impressive speed, makes it a significant development. The presenter hopes for Trey to make the model more generally available for further testing.
The video concludes with a call for viewer engagement and subscription.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

99% Follow Goals, Only 1% Do this
Him-eesh Madaan

Why Does This Guy Appear In Kids Videos?
sphynx

NVIDIA Monopoly is DEAD | OPEN-SOURCE Chips Are HERE!
Hefty LLM

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

¿Trabajas en Oficina? EL ERROR que comete el 99% con Julieta Manzano | Martha Debayle
Martha Debayle

How East India Company Captured India | Nitish Rajput | Hindi
Nitish Rajput @

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED