THE SUMMARYAI-generated
Key Concepts:
- Deepseek V3.1 Terminus: An improved version of Deepseek, excelling in tool calling and function calling, without Chinese token issues.
- Non-Reasoning Model vs. Reasoning Model: Two variants of Deepseek V3.1, with the non-reasoning model performing better in coding tasks.
- Agentic vs. Non-Agentic Tasks: Agentic tasks involve more complex, autonomous problem-solving, while non-agentic tasks are simpler and more direct.
- Kilo Code and Rue: Platforms for AI coding where Deepseek V3.1 can be used.
- Requesty: A tool that allows direct access to Deepseek's official API, bypassing issues with Open Router's implementation.
- Open Router: A platform that, according to the speaker, may not implement Deepseek V3.1 correctly, forcing the use of the reasoning model.
- Biterover: A memory layer MCP (Memory Context Provider) that allows AI coders like DeepSeek to save and manage memories for better performance.
Deepseek V3.1 Terminus Overview:
- Launch and Improvements: Deepseek launched the V3.1 Terminus model, an improved version of Deepseek, with released weights and resources. Key improvements include better tool and function calling and the elimination of Chinese token issues, enhancing reliability.
- "Terminus" Name: The name "Terminus" suggests it might be the final model in the V3 lineup, with V4 coming next.
- Coding Performance: The model is suitable for use in platforms like Kilo Code or Rue as an AI coder, with significantly reduced errors.
Testing and Benchmarks:
- Non-Agentic Benchmark: Deepseek V3.1 scores third position on non-agentic benchmark tasks.
- Reasoning Model Issues: The reasoning model performs worse than previous versions, struggling with math questions and failing to complete reasoning tasks even with multiple attempts.
- Non-Reasoning Model Success: The non-reasoning model is a major improvement.
- Specific Task Results:
- Floor plan generation is good, but navigation is problematic.
- SVG Panda with a burger is rendered adequately.
- Pokeball generation fails.
- Chess board with autoplay works well, following most rules and providing logs.
- Kandinsky style Minecraft generation is good.
- Butterfly in a garden is acceptable, despite minor issues.
- CLI tool in Rust works fine.
- Blender script performance is not as good.
- General Task Performance: The model performs well in general tasks without unnecessary reasoning, unlike previous versions that increased costs and created a bad experience by reasoning within answers even when not required.
Agentic Four Question Test with Kilo Code:
- Open Router Issue: Open Router and platforms like Rue and Kilo Code force the use of the reasoning model with Terminus, leading to a poor experience.
- Requesty Solution: Using Deepseek's API directly or via Requesty (which has a Deepseek chat endpoint) provides a better experience by allowing the use of the non-reasoning model.
- Performance Difference: The difference in performance between the reasoning and non-reasoning models is significant.
- Movie Tracker App: The model created a movie tracker app with no diffit failures or terminal errors, a rare achievement for open models. The app features a human-like design with search, calendar, and profile views, including a "clear all data" option.
- TUI Calculator in Go: The model successfully created a functional TUI calculator in Go.
- Godo Game Implementation: The model implemented a life bar affected by jump and a step counter in Godo, a task few open models can handle. The implementation was usable and done in one shot.
- Open Code Task Failure: The model failed at a large open code task due to context limit issues.
- Cost and Ranking: The model costs 50 cents total and beats Codec, ranking second after Code Buff.
Model Preference and Implementation Issues:
- Best Open Coding Model: The speaker considers Deepseek V3.1 (non-reasoning) the best open coding model, surpassing GLM and Kimmy.
- Open Router/Tool Implementation: The speaker expresses frustration with the incorrect implementation of the model provider routing in Open Router and other tools.
- Requesty and Official Endpoint Recommendation: Recommends using Requesty's Deepseek chat endpoint or the official Deepseek endpoint to avoid issues.
Context Window and Memory Management:
- Context Window Limit: The model has a 128,000 context window, which can be limiting.
- Biterover Solution: Biterover, a memory layer MCP, allows Deepseek to save and manage memories, improving performance in long-running tasks. It also enables team sharing of memories.
Conclusion:
Deepseek V3.1 Terminus, particularly the non-reasoning model accessed via Requesty or the official Deepseek API, is presented as the best open-source coding model currently available. It excels in various coding tasks, producing human-like and functional applications with minimal errors. The speaker emphasizes the importance of proper implementation to avoid the inferior reasoning model and suggests using Biterover to extend the model's memory capabilities.
AI summaries can miss context or contain errors. Check important details against the original video.