Stanford Robotics Seminar ENGR319 | Autumn 2025 | Embodied Foundation Models

Unknown AuthorAbout 8 min readNov 25, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Embodied Foundation Models: The overarching theme of the research, aiming to develop AI systems that can perceive, understand, and act in the physical world.
  • Mobility and Autonomy: Two key axes for categorizing robotic capabilities, with the ultimate goal being high levels of both.
  • Reinforcement Learning (RL): A machine learning paradigm used to train policies for robotic control, particularly for locomotion.
  • Sim-to-Real Transfer: The process of transferring policies trained in simulation to real-world robotic systems.
  • Diversity in Training Data: The importance of using varied and realistic environments and assets for training AI models to improve generalization.
  • Visual Locomotion: Training robots to navigate using visual input (RGB images) rather than solely geometric data.
  • Adaptation: Enabling robots to learn and adjust their behavior in real-time based on environmental feedback.
  • Abstraction: Leveraging higher-level representations, such as language, to guide robot behavior.
  • Vision-Language Models (VLMs): Models that combine visual and linguistic understanding for tasks like navigation.
  • Generalization: The ability of a robot to perform tasks in novel and unseen environments.
  • Tiamat Project: A large-scale research initiative focused on transferring imprecise and abstract models to autonomous technologies.

Embodied Foundation Models: Advancing Robotic Capabilities

This talk explores advancements in robotic capabilities, focusing on the development of "embodied foundation models" that aim for high levels of both mobility and autonomy. The speaker categorizes existing robotic systems along these two axes, noting that industrial robot arms have low mobility and autonomy, while self-driving cars and warehouse robots exhibit increased autonomy but are limited to specific environments. The ultimate goal is to achieve robots with high mobility and autonomy for applications like last-mile delivery, space robotics, search and rescue, and environmental monitoring.

Understanding Autonomy and Navigation

Autonomy is defined as the ability to independently perceive, understand, and act. For robots, this involves sensing the environment, processing this information, and selecting appropriate motor actions. In navigation, autonomy means achieving a goal, such as finding a safe path from a starting point to a destination.

Evolution of Quadrupedal Robot Locomotion

The field of quadrupedal robot locomotion has seen significant progress over the past decade. Early systems relied on model predictive control for static gaits. More recent advancements utilize reinforcement learning (RL) policies, enabling robots to perform dynamic maneuvers like jumping over boxes. Navigating complex terrains requires reasoning about both the environment (e.g., terrain height) and the robot's hardware capabilities. The controller and software are crucial for effective navigation.

Primer on RL Locomotion

Reinforcement learning for locomotion involves training a policy ($\pi$) to output motor commands (e.g., target joint positions). These commands are fed into low-level PD controllers that apply torque to the robot's actuators. A simulator steps the environment forward based on these torques, and a reward function guides the RL loop to optimize the policy. Sim-to-real transfer involves removing the simulator and deploying the learned policy on a real robot. Key works in this area have laid the foundation for current RL locomotion control. Training often involves simulating thousands of robots in parallel to accelerate learning.

Building Towards Embodied Foundation Models: Three Research Aspects

The research presented focuses on three key aspects of building embodied foundation models:

  1. Diversity: Increasing the variety of training domains and data.
  2. Abstraction: Leveraging higher-level representations for learning.
  3. Generalization: Enabling policies to perform in novel environments.

1. Diversity in Training Paradigms

A core argument is that AI systems often fail outside their training domains. To address this, the research aims to understand when and why quadripedal robots fail. This led to the collection of a large, openly available dataset by deploying robots in diverse environments, including caves and high-altitude mountains in Switzerland.

Data Collection and Processing:

  • Robots are equipped with sensor payloads (LiDAR, geometric sensors, cameras) for environmental sensing.
  • Collected sensor data is aggregated and fused over time and space to create elevation maps.
  • State estimation and classical algorithms are validated across diverse environments.
  • Volumetric mapping frameworks are used to build 3D scene representations for better navigation.
  • Gaussian splats are created from recorded images to generate realistic 3D environments.
  • This dataset is available for research and has been adopted by numerous institutions.

GausChim: Visual Locomotion from Pixels:

  • Problem: Traditional simulation environments for locomotion training are oversimplified, relying on heuristics and geometric inputs, leading to poor sim-to-real transfer. Rendering visual information is also challenging.
  • Solution: The GausChim project addresses these limitations by equipping robots in simulation with front-facing cameras and training them to navigate directly from RGB vision.
  • Requirements: This requires a fast visual rendering engine, accurate physics simulation, and a diverse set of visual assets.
  • Asset Generation: Instead of handcrafted heuristics, diverse assets are generated using:
    • Scanning real-world scenes with smartphones.
    • Utilizing existing datasets (e.g., the collected quadripedal data).
    • Employing video generation models to create 3D scenes.
  • Methodology:
    • VTT (View Transformation Transformer): Provides image intrinsics, extrinsics, point clouds, and spatial scene understanding.
    • Gaussian Splats: Trained based on visual information and camera poses.
    • Neural Surface Reconstruction: Obtains accurate scene geometry.
    • Physics Simulation: Combined with reconstructed geometry to create realistic simulation environments.
  • Outcome: GausChim enables training policies that can walk up and down stairs purely from RGB vision, demonstrating effective sim-to-real transfer for visual locomotion. The project provides a large asset library of 3D Gaussian splats for research.

Adaptive Policies for Traversability:

  • Alternative to Diversity: Instead of solely increasing training domain diversity, the research explores making policies more adaptive.
  • Traversability Score: A metric is developed to estimate the ease or difficulty of traversing an environment. This is derived by measuring the difference between the robot's commanded velocity and its achieved velocity.
  • Real-time Adaptation: A classifier is trained to predict traversability scores directly from images in real-time during deployment.
    • Initially, a randomly initialized network shows uncertainty (bluish color).
    • During operator-guided interaction, the network learns and updates in real-time.
    • After a few minutes of adaptation, the operator can disengage, and the robot navigates autonomously using vision-based traversability.
  • Benefits: This approach allows robots to adapt to novel surfaces (e.g., an orange floor) quickly. It also highlights the importance of vision for navigation, especially in scenarios where geometric sensors fail (e.g., detecting glass).

2. Leveraging Abstraction: Vision-Language Models for Navigation

Language models exhibit strong abstraction capabilities, enabling easy encoding of instructions and environmental understanding. The research investigates the performance of vision-language models (VLMs) in predicting navigation paths.

Navigation as Visual Question Answering (VQA):

  • Task: Given a task instruction (e.g., "go downstairs"), the VLM predicts a trace (a sequence of points in image space) to achieve the task.
  • Assumptions:
    • Most navigation tasks can be solved if a correct path can be drawn on an image.
    • Existing robot systems are capable of following such image-space paths.
  • Contributions:
    • Formulating navigation as a VQA task.
    • Collecting 1,000 navigation scenarios and labeling over 3,000 traces.
    • Developing a metric to evaluate the alignment of predicted traces with human preferences.
  • Benchmarking: A comprehensive study compared human performance with state-of-the-art VLMs.
    • Humans significantly outperform VLMs in generating navigation traces.
    • A major failure mode for VLMs is identifying the goal and determining where to go.
    • VLMs perform uniformly poorly across different embodiments and navigation challenges.
  • Conclusion: Significant work is needed to improve VLMs' understanding for navigation tasks.

3. Generalization and Open-World Tasks: The Tiamat Project

The Tiamat project aims to transfer imprecise and abstract models to autonomous technologies, focusing on generalization for open-world tasks.

Tiamat Architecture:

  • Core Idea: Perfect simulation is unattainable; therefore, abstractions are needed for sim-to-real transfer.
  • Challenge: Performing complex, open-world navigation tasks based on natural language instructions (e.g., "follow the roads to the site, locate a cluster of buildings, systematically search for them").
  • Methodology:
    • Fine-tuning a language model with inputs including:
      • Task description (natural language).
      • Environment description (robot's surroundings).
      • Embodiment tokens (latent states of RL locomotion policies).
      • Additional signals (e.g., time remaining).
    • The LLM predicts which behavior to use (e.g., navigation agent, exploration policy) and provides goal commands.
    • The system is adapted and fine-tuned in a closed loop.
  • Goal: To enable robots to generalize to a wide variety of natural language commands and perform open-world general-purpose tasks.

Synthesis and Conclusion

The research presented makes significant strides in advancing robotic capabilities by focusing on embodied foundation models. Key takeaways include:

  • Diversity is crucial: Increasing the diversity of training data, simulation environments, and assets is essential for robust robot performance.
  • Visual perception is key: Training robots to navigate using RGB vision offers a powerful alternative to purely geometric sensing.
  • Adaptation enhances robustness: Real-time adaptation allows robots to learn and adjust to novel environments and situations.
  • Abstraction aids understanding: Leveraging language and vision-language models can provide higher-level guidance for robot behavior.
  • Generalization is the ultimate goal: Projects like Tiamat aim to equip robots with the ability to perform complex, open-world tasks based on natural language instructions.

The speaker concludes by emphasizing the ongoing challenges and the exciting future of embodied AI research.

Discussion on Manipulation

When asked about applying these ideas to manipulation with robot fingers, the speaker highlights key challenges compared to locomotion:

  • Force Accuracy: Manipulation requires precise control of forces applied to objects, which is more critical than in locomotion where forces are less sensitive. Simulation accuracy for forces is paramount for manipulation.
  • Solution Space: Locomotion, especially with quadrupedal robots, offers a larger solution space and redundancy, allowing for more flexibility. Manipulation tasks, like picking up an egg, have an extremely small and precise solution space, demanding exact forces.
  • Sim-to-Real for Manipulation: It remains an open question whether simulation is the optimal approach for sim-to-real transfer in manipulation. Imitation learning has shown promise, but RL fine-tuning in the real world might be necessary.

Regarding the RL primer diagram and the choice of position control over force control:

  • Initialization: Starting with position control provides a stable initial state (standing) for the robot, making the RL search problem easier.
  • Discovery: RL training involves applying noise to discover skills. With force control, the robot would need to actively compensate for gravity and maintain balance from the outset, making the search problem significantly harder.
  • Performance: Attempts to use force control for RL locomotion have consistently shown worse performance compared to position control-based setups.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.