Stanford Robotics Seminar ENGR319 | Winter 2026 | Autonomous Navigation in Outdoor Environments

By Stanford Online

Share:

Key Concepts

  • Robot Navigation: Planning and autonomous movement of a robot to a goal.
  • Traversibility: The ability of a robot to move across different terrains.
  • Gaussian Splatting: A method for rendering the environment around a robot with semantic information.
  • Vision-Language Models (VLMs): Models that combine visual and textual information for understanding and decision-making.
  • Social Navigation: Robot navigation that considers pedestrian safety and comfort.
  • Compliant Control: Robot control that allows for safe and adaptable interaction with the environment and humans.
  • Grounded Reasoning: Using physical understanding and perception to inform decision-making in AI models.
  • Humanoid Robotics: Development of robots with human-like form and capabilities.
  • Large Language Models (LLMs): Powerful models capable of understanding and generating human language.
  • Low-Rank Adaptation (LoRA): A technique for efficiently fine-tuning large language models.

Traversibility Analysis and Robot Navigation in Outdoor Environments

The research initially focuses on robot navigation, particularly in challenging outdoor environments. Robot navigation is defined as the process of a robot planning a path and autonomously moving to a goal, with significant implications for logistics, delivery, and warehouse management. A key challenge is traversibility – determining which areas a robot can navigate. This depends on the robot’s design; flat surfaces suffice for wheeled robots, while legged robots can handle curbs and stairs. Vegetation presents a variable obstacle, traversable by larger, heavier robots but not smaller ones.

To address this, the researchers developed a system that generates potential trajectories using an autoencoder-decoder mechanism. This involves projecting Gaussian distributions onto orthogonal axes to create multiple trajectory candidates. These candidates are then evaluated based on distance to the robot and selected using Vision-Language Models (VLMs). A pipeline was created where VLMs choose the best trajectory, demonstrating improved performance compared to purely VLM-based or learning-based approaches. However, the system still encounters issues in out-of-distribution scenarios, specifically unexplained failures in certain terrains (TJS). Extensive training was found to mitigate these issues.

Navigation Dataset and Environmental Representation

To improve generalization, a new navigation dataset was created, encompassing 10 campuses and 11 hours of ROS bag data. This dataset includes global traversibility maps with five categories, correlated with robot positions and sensor data (3D LiDAR, RGB camera, GPS, odometer).

The researchers then moved beyond simple traversibility mapping to more detailed environmental representation. They leverage Gaussian Splatting to render the environment with semantic information. To improve geometric accuracy, LiDAR information is fused using Euclidean Signed Distance Fields (ESDF). This combined map is used for path planning. However, semantic information alone isn’t sufficient; the robot needs to understand the physical properties of the terrain (pliability of grass, traversability of bushes).

To address this, the system jointly estimates semantic material types (grass, trees, buildings) using a Gaussian posterior and physical properties (friction, hardness, stiffness, density) using a Dirichlet Categorical Posterior. This results in a local map where red indicates non-traversable areas and colored points represent Gaussian splats. This map enables the robot to navigate unstructured vegetation effectively.

Socially Compliant Navigation

The research extends to social navigation, focusing on safe and comfortable interaction with pedestrians and adherence to traffic rules. This is broken down into three steps: perception (identifying pedestrians and their movements), prediction (forecasting pedestrian motion and potential interactions), and action (planning robot movement based on social and traffic context).

VLMs are leveraged to understand social and traffic cues. A new dataset, SNE (Social Navigation with Language context), was created for training and evaluation. Evaluations of ChatGPT, Gemini, and a fine-tuned LLaVA model (using Low-Rank Adaptation - LoRA) showed that the fine-tuned LLaVA generally outperformed the others in predicting pedestrian behavior. Gemini also performed better than ChatGPT.

Integrated Navigation Stack and Real-World Performance

The individual components (traversibility analysis, social navigation) are integrated into a complete navigation stack, combined with GPS routing and low-level motion planning. This stack has been tested on various robot platforms (legged and wheeled) in diverse urban and rural scenarios. The robots demonstrate the ability to navigate pavements, sidewalks, cross streets following traffic rules, and interact compliantly with pedestrians.

Extending Navigation to Companion Robots for Older Adults

The current research direction focuses on extending the navigation stack to companion robots designed to assist older adults. The demographic shift towards an aging population (projected to be over 20% by 2030) motivates this work. Companion robots are envisioned as enhancements to vision, hearing, and mobility, enabling independent and healthy living.

Two key research areas are identified:

  1. Navigation Assistance: Creating a navigation stack that acts as “eyes and ears,” alerting users to hazardous terrains and traffic conditions. This requires a real-time, accurate world model.
  2. Behavior Analysis: Monitoring older adults’ movements to identify potential health risks (e.g., falls) without restricting their independence. This involves using robots to analyze motion and correlate it with potential diseases.

Leveraging LLMs for Humanoid Robot Interaction (New Research)

A significant portion of the presentation shifts to research on utilizing LLMs for humanoid robot interaction. The goal is to create robots that can understand human intentions, reason about context, and perform tasks safely and effectively. The speaker highlights the need to move beyond simply replicating human motions to enabling robots to adapt to unexpected situations and interact naturally with people.

Key approaches include:

  • Gentle Humanoid: Developing compliant control strategies that allow robots to safely interact with humans, adapting to external forces and maintaining stability. This involves modeling interaction forces and using a reward function during training to prioritize safety.
  • ChatPost & ChatWoman: Utilizing LLMs to understand human behavior and generate appropriate 3D poses and motions. This involves projecting 3D pose information into the language space, allowing the LLM to reason about human actions and predict future movements. The system can interpret text descriptions and generate corresponding robot motions.
  • Reasoning and Planning: Developing a system where the LLM acts as an “agent,” reading research papers and utilizing tools to solve problems related to human interaction. This aims to create a more intelligent and adaptable robot that can handle complex scenarios.

Future Directions and Collaboration

The speaker concludes by outlining future research directions, including improving real-time performance, handling dynamic human environments, and incorporating emotional intelligence into robot interactions. They emphasize the need for collaboration and invite interested researchers to contact them. Specific areas for future work include:

  • Developing more accurate biomechanical models of humans, particularly those with mobility constraints.
  • Integrating multimodal expressions (voice, haptics, facial expressions) to enhance communication.
  • Utilizing humanoid robots to collect data for improving AI models.

Notable Quotes:

  • “Ultimately by offloading the tedious and risky tasks to autonomous agents we can release ourself and we can free oursel to focus more on creative and high value works.”
  • “We all want to stay independent, engaged, and vibrant. And this is where our companion robots can step in.”
  • “We need to create a real time and also accurate word model to handle all the emergencies and hazardous terrains.”

This summary provides a detailed overview of the presented research, preserving the technical precision and language of the original transcript. It aims to be a comprehensive resource for understanding the advancements in robot navigation and the exciting potential of companion robots for improving the lives of older adults.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video