DeepMind’s New AI Found A Strange New Way To Think

Two Minute PapersAbout 3 min readJun 5, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • AlphaProof Nexus: A DeepMind AI system designed to solve complex, long-standing mathematical problems.
  • Lean: A formal mathematical language used to verify the correctness of proofs, preventing AI "hallucinations."
  • Tournament-based Iteration: A methodology where AI-generated solutions compete against each other, receiving ELO scores to iteratively improve upon the best-performing "bad" solutions.
  • Reliable Systems from Unreliable Parts: The core philosophy that an unreliable AI can produce a verified, correct result if placed within a rigorous, iterative "harness" or loop.
  • First Law of Papers: A heuristic suggesting that one should evaluate AI progress based on future potential rather than current limitations.

1. Main Topics and Performance

DeepMind’s AlphaProof Nexus recently attempted to solve approximately 350 of the legendary Paul Erdős’s open mathematical problems. While the system had a 95.7% failure rate, it successfully solved nine problems that had remained unsolved for decades.

  • Cost Efficiency: Each problem was solved at a cost of only a few hundred dollars.
  • Significance: Critics argue the AI did not perform "fundamentally new" tasks, but the author contends this ignores the rapid trajectory of AI development—moving from basic arithmetic (4 years ago) to high school competition problems (2 years ago) to Olympiad gold medals (1 year ago), and now to decades-old unsolved problems.

2. Methodology: The Tournament Framework

The breakthrough lies in the "harness" or loop surrounding the AI, rather than just the intelligence of the model itself. The process follows these steps:

  1. Formalization: A mathematician translates the problem into Lean, ensuring the proof can be programmatically verified.
  2. Initial Attempt: The AI agent attempts to solve the problem.
  3. Evaluation: A "judge" AI reviews the solution. If it fails, it provides feedback on why the solution was inadequate.
  4. Tournament Loop: A cheaper "judge" AI compares two previous (incorrect) solutions and selects the one that is slightly better.
  5. ELO Scoring: Each solution is assigned an ELO score (named after Arpad Elo). The system then iterates, using the highest-scoring "bad" solution as the starting point for the next round.
  6. Verification: This process repeats until the validator confirms the proof is mathematically sound.

3. Key Arguments and Perspectives

  • Intelligence in the Loop: The author argues that the future of AI is not just about making models "smarter," but about designing tighter, more effective harnesses. By allowing an unreliable model to fail thousands of times against a rigid, truthful judge, the system eventually converges on a correct answer.
  • The "First Law of Papers": Do not judge AI by its current state, but by the trajectory of its development. The rapid progression in mathematical problem-solving suggests that current limitations will likely be overcome in subsequent iterations.

4. Limitations and Research Findings

  • Selection Bias: The 350 problems tested were likely a subset chosen for their ease of formalization, rather than a random sample of all 1,200+ Erdős problems.
  • Model Size Dependency: Smaller AI models failed to solve any problems, confirming that a "beefy" or frontier-level AI system is still required at the core of the process.
  • Resource Allocation: An open question remains regarding the trade-off between using larger models with fewer tournament rounds versus smaller models with more rounds, assuming equal computational costs.

5. Synthesis and Conclusion

The success of AlphaProof Nexus marks a paradigm shift in AI research. By moving away from the expectation that a single model must be perfect, researchers have demonstrated that reliable systems can be constructed from unreliable parts through iterative, tournament-based validation. While the system is not yet perfect and relies on specific, formalizable problem sets, the ability to solve 56-year-old mathematical mysteries for a few hundred dollars represents a stunning leap in computational capability. The intelligence of the future lies in the architecture of the loop, not just the model itself.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.