The LLM Revolution Is Over. The Physical AI Revolution Is Coming Fast
By Forbes
Key Concepts
- AGI (Artificial General Intelligence): The speaker rejects the term, arguing human intelligence isn’t truly “general” and focusing on surpassing human capabilities is more accurate.
- LLMs (Large Language Models): Current dominant AI paradigm, but seen as limited due to lack of “world models” and inability to predict consequences.
- World Models: Crucial for intelligent behavior; systems that can anticipate the consequences of actions and plan accordingly. Missing component in current LLMs.
- Generative vs. Non-Generative Architectures: Generative architectures (like those used in LLMs) are ill-suited for understanding the complexities of the real world. Non-generative architectures, like Japa (Joint Embedding Predictive Architecture), are being explored at AMI.
- Objectived-Driven AI: A blueprint for AI systems focused on achieving a defined objective with built-in safety guardrails.
- Open vs. Closed AI: Advocates for open-source AI to accelerate progress, particularly outside of China and the US, and to ensure a diverse and representative AI ecosystem.
- Physical AI: The next revolution in AI, focusing on understanding and interacting with the physical world through sensory data (video, sensor data).
The Path to Advanced AI & Limitations of Current Approaches
The speaker begins by dismissing the term “AGI,” arguing that human intelligence isn’t “general” enough to serve as a useful benchmark. He believes surpassing human intelligence is inevitable, but not imminent, requiring “a few conceptual breakthroughs.” He emphasizes that simply scaling up current LLM paradigms won’t achieve this.
A fundamental flaw in current AI, particularly LLMs, is the lack of “world models.” LLMs excel at predicting the next word, but lack the ability to anticipate the consequences of actions or plan effectively. He illustrates this with the contrast between a 10-year-old solving a new task intuitively versus the massive training data required for even limited autonomous driving capabilities. “If you want intelligent behavior, you need a system to be able to anticipate and what's going to happen in the world and and also predict the consequences of its actions.”
The Importance of Understanding the Real World
The speaker stresses that “the real world is way more complicated than the world of language.” While language appears to be the pinnacle of human intelligence, predicting the next word is comparatively simple. True intelligence stems from understanding the physical and social world.
Current generative architectures, used in LLMs, are unsuitable for processing the “messy,” high-dimensional, continuous, and noisy data of the real world. The next AI revolution will be driven by systems capable of understanding this data – video, sensor data – and building predictive models of their environment. These systems must also be controllable and safe.
Meta AI Research & The Rise of Open Source
During his 12 years leading AI research at Meta, the speaker identifies the openness of AI research as the biggest factor in its rapid progress. Sharing research papers and open-sourcing code fostered collaboration and accelerated innovation. However, he expresses concern about the recent trend towards closed AI research within industry labs (Anthropic, Google, and a shift at Meta), which he believes will slow progress, particularly in the West. He notes that the best open-source models are currently emerging from China. “It’s not any particular contribution [like transformers]. It’s the fact that AI research was open.”
Advanced Machine Intelligence (AMI) & Joint Embedding Predictive Architecture (Japa)
AMI is focused on building a new generation of AI systems based on world models, learning from sensory data (video, physical interaction, spatial data) rather than language alone. The core of this approach is a non-generative architecture called Japa.
Japa trains systems to understand and predict video content, even recognizing impossible scenarios. The goal is to generalize this methodology to any data modality, enabling the creation of phenomenological models of complex systems – industrial processes, chemical plants, even living cells. He contrasts this with the limitations of overly accurate simulations, arguing that abstract representations are crucial for prediction. “We have systems now that we can train completely self-supervised on unlabeled videos and those systems understand video represent it really well can predict missing parts in a video and they also have acquired a certain sense of common sense.”
Open vs. Closed AI: A Platform Perspective & Global Implications
The speaker views AI as becoming a platform, historically driven by open-source principles (like the internet). He argues that proprietary AI systems will be less widely adopted. He advocates for a global consortium to train an open-source LLM as a repository of all human knowledge, emphasizing the need for multilingual and culturally diverse data.
He warns that if AI is controlled by a few companies in the US or China, it poses a significant threat to democracy, cultural diversity, and linguistic diversity. “The biggest risk of AI is that in the near future where our entire digital diet will be mediated by AI systems if those AI systems come from a handful of proprietary uh companies on the west coast of the US or China, we're in big trouble.”
AI Risks & Alignment: Beyond Apocalyptic Narratives
The speaker dismisses “apocalyptic” AI narratives as distractions. He identifies the concentration of power among a few companies/governments as the most pressing risk, followed by human misuse of AI. He believes economic displacement will be less severe than predicted, as the rate of technology adoption is limited by people’s ability to learn new skills.
He reframes the “alignment” problem, arguing that it’s not about controlling LLMs, but about designing fundamentally different AI architectures (like objectived-driven AI) with built-in safety guardrails. “So if you try to project if you imagine that future AI systems that have humanlike intelligence will be LLMs, which of course is not going to happen, you say, "Oh my god, that's going to be dangerous." It's the wrong approach.”
Preparing for an AI-Rich Future
The speaker advises students to focus on fundamental knowledge with a long shelf life (e.g., quantum mechanics over mobile app programming) and to be prepared for continuous learning and career changes. He emphasizes that AI will augment human intelligence, assisting us in decision-making and amplifying our capabilities. He envisions a future where AI assistants are integrated into our daily lives through wearable devices.
Looking Ahead to 2035
In 2035, success would involve AI systems that understand the physical world and potentially reach human-level intelligence (or surpass it in specific domains). He anticipates conceptual breakthroughs, not a single “secret” to AGI, and emphasizes that progress will be incremental and often overlooked initially. He believes AI will become a ubiquitous tool, assisting us in all aspects of life, and that our relationship with advanced AI will resemble that between leaders and their staff – leveraging superior intelligence to achieve goals. He predicts that the next five years will see even faster change than the previous five, driven by breakthroughs in research that are currently underappreciated.
Notable Quotes
- “Calling human level AI AGI is a misnomer.”
- “If you want intelligent behavior, you need a system to be able to anticipate and what's going to happen in the world and and also predict the consequences of its actions.”
- “The real world is way more complicated than the world of language.”
- “AI is fast becoming a platform and historically platforms have always become open source.”
- “The biggest risk of AI is that in the near future where our entire digital diet will be mediated by AI systems if those AI systems come from a handful of proprietary uh companies on the west coast of the US or China, we're in big trouble.”
Synthesis/Conclusion
The speaker presents a nuanced and critical perspective on the current state of AI. He dismisses hype surrounding AGI and LLMs, advocating for a shift towards AI systems grounded in understanding the physical world through robust “world models.” He champions open-source AI as crucial for accelerating progress and ensuring a diverse and equitable future. His vision for AMI represents a departure from the dominant paradigm, focusing on non-generative architectures and learning from sensory data. Ultimately, he believes the next decade will bring significant advancements, but emphasizes the importance of focusing on fundamental research, open collaboration, and responsible development to harness the full potential of AI.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Stanford CS153 Frontier Systems | Building the Frontier Ecosystem
Stanford Online

'Things are going to be okay, in Canada and the U.S.': Thorne
BNN Bloomberg

I'M OUT: The $11 Trillion AI Bubble is Breaking!
Steven Van Metre

South Korea bets big on AI with nearly a trillion dollars of investment • FRANCE 24 English
FRANCE 24 English

The Bubble is Bursting... (Emergency Update)
Bravos Research

The AI Bubble Just Ended - Without Popping
Heresy Financial

AI Market Volatility, Europe Heat Wave, Venezuela Quakes Damage | Bloomberg This Weekend: June 27
Bloomberg Television