Stanford Robotics Seminar ENGR319 | Autumn 2025 | The Graph Physical AI Approach

By Unknown Author

Share:

Key Concepts

  • Physical AI: AI applied to robots operating in the real, physical world, distinct from digital AI.
  • Foundation Models: Large, general-purpose AI models that can be adapted for various downstream tasks.
  • Data-Driven AI: AI models trained on large datasets to learn patterns and make predictions.
  • Physics-Informed AI: AI models that incorporate physical principles (kinematics, dynamics, etc.) into their architecture and training.
  • Cross-Embodiment Learning: The ability of an AI model to transfer knowledge learned from one physical embodiment (e.g., a specific robot) to another.
  • Agentic Framework: A system where specialized AI agents interact to perform complex tasks, enabling modularity and interpretability.
  • VLA (Vision, Language, Action) Models: AI models that integrate understanding from visual input, natural language instructions, and desired robot actions.
  • Kinematics: The study of motion without considering the forces that cause it, focusing on the geometry of movement.
  • Soft Physics: Encoding of physical concepts that are not purely kinematic, such as soft contact, stacking, and gravity.
  • Real-time Performance: The ability of an AI model to process information and respond within strict time constraints, crucial for robotics.
  • Flywheel Effect: A continuous cycle of data collection, product development, and release, leading to rapid improvement.

Summary

The talk explores the integration of AI, particularly foundation models, into robotics for real-world applications, emphasizing the critical role of physics in achieving true physical intelligence.

The Evolution of AI in Robotics

The speaker begins by highlighting the generational upgrade in AI, contrasting the vast data and knowledge available for digital AI with the nascent stage of physical AI for robots. While digital AI benefits from decades of human data, robots are still catching up. The success of autonomous driving (Waymo) is presented as an early proof point, but the focus shifts to enterprise robots as the next frontier for mass adoption due to their integration into environments already accustomed to machines.

Real-World Applications and Market Opportunities

The speaker identifies key enterprise robot applications: dexterous assembly, material handling, construction, agriculture, and distribution. A compelling example is the food industry, where 70% of food waste occurs not due to production but distribution failures. Similarly, during COVID-19, essential medicines were lost due to distribution bottlenecks. These fundamental problems represent significant market opportunities, estimated to be in the hundreds of billions of dollars. The speaker emphasizes that solving these problems creates value rather than engaging in a zero-sum game.

The Foundation of Data-Driven AI and its Limitations

The core principles of data-driven AI are outlined: defining input/output, acquiring training data, and specifying a loss function for optimization. Early work on single-image depth estimation is cited as an example where framing the problem correctly and using shallow models like Markov Random Fields, combined with a few thousand data points, unlocked new capabilities. The evolution to deeper architectures like Convolutional Neural Networks (CNNs) further democratized applications like OCR and image detection, becoming a staple in robotic perception.

However, the speaker identifies a significant limitation of purely data-driven approaches in robotics: the "data problem." Collecting sufficient, diverse data for physical AI is prohibitively expensive and time-consuming. Examples like Tesla's self-driving and Waymo's extensive data collection efforts, costing billions, illustrate this challenge. The speaker argues that this approach is not sustainable for diverse robotic applications like underwater, space, mining, or household robots.

The Need for Physics in Physical AI

The speaker recounts personal experiences, such as attempting to teach a robot to cook, where collecting vast amounts of data (e.g., cutting cucumbers) did not yield the desired results. This led to the realization that robots operate in the physical world, and tasks like cutting food are incredibly complex due to the nuances of physical interaction.

The speaker critiques current VLA (Vision, Language, Action) models, stating they often miss the crucial point of physical AI: the integration of physics. While these models excel at transferring knowledge across modalities (e.g., language concepts to French), they struggle with the "out-of-distribution" nature of physical data. The embodiment of humans in training data (e.g., YouTube videos) differs from robot embodiments, creating a transfer gap.

Encoding Physics for Robust Robotics

The central argument is that physics must be a foundational element of physical AI, not an afterthought. The speaker proposes incorporating kinematics, the study of motion, as a starting point. This can be achieved by representing robots and tasks as graphs, where nodes represent components (head, arms, legs) and edges represent their relationships. This graph representation can then be used to train models that understand physical constraints.

The speaker's team at Togi is building a framework that integrates physics into the AI pipeline:

  1. Embeddings/Encoders: Processing various input data (robot, haptic, 3D).
  2. The Model: A large model that minimizes a loss function, now incorporating physics.
  3. Agents: Thin interfaces that interact with the model, enabling modularity and task execution.

Key Physics Integration Techniques:

  • Kinematics: Encoding the robot's kinematic structure and its relationships. This allows for transfer learning across different robot embodiments.
  • Soft Physics: Incorporating concepts like soft contact, stacking, and gravity through small neural networks accessible to the main model.
  • Physics Neural Operators: Wrapping known physical equations within a loss function, allowing for deviations based on data.
  • Inverse Kinematics Simulator: Integrated into the model to perform reasoning at the microsecond level, significantly speeding up the reasoning loop.

Real-time Performance and GPU Acceleration

A critical challenge for physical AI is achieving real-time performance. The speaker highlights a partnership with Nvidia to leverage GPUs for model and data processing, keeping everything in GPU memory. Agents act as thin interfaces on the CPU, interacting with the GPU-based model. This is crucial for robots that require responses within 100 milliseconds to avoid getting stuck.

Case Studies and Real-World Deployments

Several examples demonstrate the application of physics-informed AI:

  • Box Folding: Robots can fold boxes with minimal training data (around 10 examples) by understanding their own kinematics and the articulated nature of objects.
  • Food Cutting: Robots can cut 25 different food items by mapping visual and haptic data within a graph architecture, allowing for reuse of the model even with different robot end-effectors.
  • Ice Cream Making (Past Example): A previous attempt was too slow, highlighting the need for real-time performance.
  • Warehouse Operations: Robots are being deployed in production environments without extensive data collection by the company itself, leveraging a "flywheel effect" where production data is used for continuous improvement.
  • Super Humanoid Robots: Assisting with heavy lifting tasks like furniture and shipping, addressing complex perception and manipulation challenges with crumpled boxes and human interaction.
  • Airport Operations: Automating tasks like baggage handling and ground operations to improve efficiency and safety, especially in adverse weather conditions where human presence is restricted.
  • Construction (Canvas Construction): Robots are used to create precise walls, requiring intricate understanding of material application and surface finishing.
  • Heavy Equipment (Farms, Mining): Operating large machines for tasks like scooping, crop collection, and targeted irrigation/insecticide application, augmenting human capabilities with high efficiency.
  • Rail Yard Logistics: Robots operate in challenging environments (rain, snow, night) to provide intelligence for specific tasks.

The Future of Robotics and AI

The speaker envisions a future where AI becomes the framework for robotics, flipping the traditional software development model. Companies will focus on building agents and improving models, rather than writing extensive code. This stratification of the robotics industry will lead to specialized companies for testing, model development, and deployment support.

The key takeaway is that intelligent use of data, with physics as the backbone, is the only way to solve the physical AI problem. This approach enables cross-embodiment learning, allowing robots to adapt to new tasks and grippers with minimal retraining, unlocking massive opportunities across various sectors.

Recommendations for Students

For students interested in VLA or robotics, the speaker recommends:

  • Embracing VLA: Utilizing VLA or similar large models.
  • Integrating Physics: Not neglecting the importance of physics in AI models.
  • Marrying Robotics and AI: Exploring techniques that combine learnings from robotics (e.g., memory concepts) with generative AI, especially for contextual real-world operations.

Regulatory Considerations

Navigating regulatory hurdles is crucial. The strategy involves operating within existing frameworks, choosing use cases with less stringent regulations (e.g., gate areas at airports vs. runways), and leveraging concepts from established industrial automation.

Robot Design and Morphology

The presented methods can inform future robot design by simulating and optimizing sensor placement and morphological configurations, accelerating the end-to-end robot design process.

Training and Data Efficiency

The focus is on reducing data requirements. By uploading robot kinematic structures and sensor configurations, the model is architected to have the capacity to learn physics, enabling new tasks with significantly fewer training examples (e.g., 100 examples instead of 10,000 or a million).

Sensor Salience and Event Detection

The discussion touches upon the importance of sensors detecting salient events, particularly time-oriented ones, which can be missed by vision or LiDAR alone. Research is needed to build better embeddings that capture this multi-modal temporal information and to encode fundamental physical concepts like object permanence.

Collaboration with Robotics Providers

When working with existing robotics providers (e.g., Dexterity), the value-add lies in augmenting or replacing specific modules with physics-informed agents, gradually unlocking more use cases and enabling smarter communication between agents.

New Trends in Robotics

The speaker is excited about the stratification of the robotics industry, leading to specialized service companies. This allows application builders to focus on their core tasks without building entire stacks from scratch. On the research side, exciting directions include understanding spatial-temporal sensor data, causality, and event detection in robotics, augmenting current VLA models with real-world physical information.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video