DeepSeek: What Actually Matters (for the everyday user)

Jeff SuAbout 6 min readMar 25, 2025Watch original
THE SUMMARYAI-generated

DeepSeek: Busting the Top 10 Myths

Key Concepts: DeepSeek, Model Efficiency vs. Performance, Export Controls, Hopper GPUs (H100, H800), Model Distillation, Inference, Edge Inference, Javon's Paradox, Chain of Thought, Open Source Models.

1. Myth #1: DeepSeek Built Their Model for Just $5.6 Million

  • Inaccuracy: The $5.6 million figure only represents the final training run cost.
  • Omitted Costs: It excludes significant infrastructure expenses, such as the reported 50,000 Nvidia Hopper GPUs, estimated to be worth around $1 billion.
  • Analogy: Comparing it to the manufacturing cost of an iPhone versus the total R&D and other associated costs.
  • Conclusion: The true investment scale is significantly larger than $5.6 million.

2. Myth #2: DeepSeek Must Have Broken the Rules

  • Innovation within Constraints: DeepSeek's success is attributed to innovating within the limitations of export controls, specifically the use of less powerful H800 GPUs.
  • Optimization: They optimized their model architecture to compensate for the H800's limitations.
  • H100 vs. H800: H100 GPUs are more powerful but were not allowed to be sold to Chinese companies, while H800 GPUs are nerfed versions that were permitted. Both are Hopper generation GPUs.
  • Analogy: Similar to Samsung using slightly slower processors in certain regions due to licensing agreements.
  • Conclusion: DeepSeek did not break any rules; they optimized their design within the existing constraints.

3. Myth #3: DeepSeek Has Beaten OpenAI

  • Performance vs. Efficiency: There's a crucial distinction between optimizing for performance (maximizing output) and optimizing for efficiency (achieving a good output with less resources).
  • Model Comparison: DeepSeek's reasoning model R1 matches OpenAI's reasoning model 01 in performance. However, OpenAI has already released 03, which is more powerful than both 01 and R1.
  • OpenAI's Response: OpenAI released 03 mini and made it available for free users, potentially due to the pressure from DeepSeek.
  • Analogy: Comparing a flagship iPhone to a cost-effective smartphone that offers 90% of the performance at a fraction of the cost.
  • Conclusion: DeepSeek leads in efficiency, but not in overall capability.

4. Myth #4: DeepSeek's Models Are Directly Comparable to All Other AI Models

  • Apple-to-Apple Comparisons: It's essential to compare models that serve similar purposes.
  • Model Equivalents: DeepSeek's base model V3 is comparable to ChatGPT's 4.0. DeepSeek's reasoning model R1, with search functionality, was initially unique until ChatGPT released 03 mini with search.
  • Analogy: Comparing a sports car to an SUV – both are vehicles, but serve different purposes.
  • Conclusion: Comparisons must be made between models with similar functionalities and intended uses.

5. Myth #5: DeepSeek R1's Visible Chain of Thought Is a Technical Breakthrough

  • UI Choice: The visible Chain of Thought in DeepSeek R1 is primarily a user interface (UI) choice, not a fundamental technical innovation.
  • Reasoning Capabilities: Both R1 and OpenAI's 01 have similar reasoning capabilities. DeepSeek chooses to display the model's thinking process to the user.
  • Analogy: Two chefs making the same dish, one in a closed kitchen and the other at a demonstration counter. The process and result are the same, but the experience differs.
  • Conclusion: The transparency is related to presentation, not core capability.

6. Myth #6: DeepSeek Built Everything From Scratch

  • Model Distillation: DeepSeek allegedly used a process called "Model Distillation," where they trained their models on the outputs of other models, such as ChatGPT.
  • Legality vs. Terms of Service: This practice is not illegal, but it violates OpenAI's terms of service.
  • Microsoft's Involvement: Microsoft is investigating DeepSeek while simultaneously adding R1 to its cloud offerings.
  • Analogy: A phone manufacturer replicating iPhone image processing by analyzing millions of iPhone photos instead of copying code.
  • Conclusion: DeepSeek likely used model distillation, which is a controversial practice.

7. Myth #7: Using DeepSeek Is Automatically Unsafe

  • Data Handling: Using the native DeepSeek app (web or mobile) sends data to and stores it in China.
  • Workarounds:
    • Use platforms like Perplexity or Venice AI to access DeepSeek's models while keeping data in the US.
    • Run DeepSeek's models locally on a desktop or laptop using applications like Ollama or LM Studio.
  • Platform Integration: More platforms are expected to integrate DeepSeek due to its cost-effectiveness.
  • Conclusion: Safety depends on data privacy concerns; workarounds exist for users who prioritize data security.

8. Myth #8: This Kills Nvidia's Business

  • Javon's Paradox: The speaker invokes Javon's Paradox, which suggests that increased efficiency leads to increased demand.
  • Increased Demand: More efficient AI like DeepSeek is likely to increase overall demand for AI solutions, potentially leading to even greater demand for Nvidia's chips.
  • Analogy: Cheaper smartphones increased demand for premium phone processors.
  • Conclusion: DeepSeek's efficiency may ultimately benefit Nvidia by expanding the AI market.

9. Myth #9: This Is Terrible for US Tech Companies

  • Potential Benefits: This development could be a win for some US tech companies.
  • Amazon: Can leverage high-quality open-source models like DeepSeek at lower costs, addressing their lack of a leading proprietary model.
  • Apple: Can leverage Apple Silicon chip advantages for "Edge Inference" (running AI models locally on devices).
  • Meta: Benefits significantly from cheaper inference, as every aspect of their business (e.g., advertising) relies on AI.
  • Inference Explanation: Inference is the application of learned knowledge to new situations.
  • Analogy: Cheaper smartphones and faster internet enabled companies like Uber and Instagram.
  • Conclusion: Cheaper AI could enable new products and services, and US tech companies are well-positioned to capitalize on this.

10. Myth #10: This Is China's Sputnik Moment in AI

  • Sputnik Comparison: Sputnik was a surprise demonstration of superior capability by the USSR.
  • DeepSeek's Transparency: DeepSeek publishes its methods openly.
  • Expected Improvements: DeepSeek achieved efficiency improvements that were already anticipated.
  • Innovation within Frameworks: They innovated within existing technological frameworks.
  • Google 2004 Comparison: More akin to Google's 2004 demonstration of building efficient infrastructure without expensive mainframes.
  • Conclusion: DeepSeek is showing that competitive results can be achieved without the most powerful chips, similar to Google's demonstration of efficient infrastructure.

Implications for Users

  1. Access to Advanced AI Features: Users now have access to powerful reasoning models (DeepSeek R1 and ChatGPT 03 mini) without paying.
  2. Data Privacy: If privacy is a concern, use platforms like Perplexity or run models locally via LM Studio or Ollama.
  3. Smart Switching: Don't switch tools solely based on trends. Only switch to DeepSeek if it provides clear advantages for specific needs.

Synthesis/Conclusion

DeepSeek's emergence is a significant achievement that highlights the potential for innovation within constraints and the importance of efficiency in AI. While the implications for US tech companies and the stock market are debatable, it's undoubtedly a win for users, providing access to advanced AI capabilities at a lower cost. However, it's crucial to understand the implications, especially regarding data privacy, and make informed decisions about adopting new tools and technologies. The speaker believes that OpenAI released 03 mini for free because of DeepSeek.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.