DeepSeek 3.1 is BETTER than Claude Sonnet 4? (FREE)

Mervin PraisonAbout 2 min readAug 22, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

Deepseek v3.1, SW bench, SWE bench multilingual, Terminal bench, Frames, Simple QA, HLE, Hybrid inference, Thinking mode, Non-thinking mode, Agentic skills, Anthropic API format, OpenAI SDK, Long context extension, Open-source weights, Token efficiency, Artificial analytics intelligent index, Claw for Sonnet thinking.

Deepseek v3.1 Release and Improvements

Deepseek v3.1 has been released, demonstrating improvements over previous versions (V3 and R1) across several benchmarks. These include SW bench, SWE bench multilingual, and Terminal bench. It also outperforms Deepseek R1 on benchmarks like Frames, Simple QA, and HLE, utilizing hybrid inference.

Thinking vs. Non-Thinking Modes

Deepseek v3.1 offers two modes: a "thinking mode" (Deepseek Reasoner) and a "non-thinking mode" (Deepseek Chat). Both modes support a context window of 128,000 tokens. The thinking mode is designed for stronger agentic skills and faster reasoning.

API and Integration

The Deepseek API supports the Anthropic API format, allowing users to integrate it into their applications using the Anthropic SDK by simply changing the base URL and API key. This is analogous to using the OpenAI SDK.

Training and Open-Source Availability

Deepseek v3.1 underwent continued pre-training on 840 billion tokens for long context extension, building upon the V3 model. The model's weights are available open-source on Hugging Face.

Pricing

When using the API directly from deepseek.com, the pricing is as follows:

  • Input: $0.01 per million tokens
  • Output: $0.56 per million tokens (with catchmiss)
  • Output: $1.68 per million tokens

The presenter notes that this pricing is comparatively cheaper than other providers.

Performance and Comparison

According to the Artificial Analytics Intelligent Index, Deepseek v3.1 is one step higher than Claw for Sonnet thinking. The presenter emphasizes the model's open-source nature and better token efficiency. A comparison with previous Deepseek versions is shown, highlighting the improvements.

Example: Ball Bouncing in Spinning Hexagon

The video showcases Deepseek v3.1's capabilities with an example of a ball bouncing in a spinning hexagon, demonstrating realistic physics and gravity simulation.

Cost Comparison

A cost comparison is presented, showing Deepseek v3.1 as cheaper than models like Qwen, GPT-4, and Gemini 1.5 Pro. However, the presenter notes that its speed is not as competitive.

Conclusion

Overall, Deepseek v3.1 is presented as a strong model, particularly considering its open-source nature and competitive pricing. The presenter invites viewers to request testing of the model and share their thoughts in the comments. He also mentions a separate video testing GPT-5, encouraging viewers to watch it.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.