Waymo's EMMA: Teaching Cars to Think - Jyh Jing Hwang, Waymo

AI EngineerAbout 4 min readJul 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Autonomous driving history, L2 vs. L4 systems, perception, prediction, planning, scaling challenges, long-tail events, foundation models, Gemini, Emma (end-to-end multimodal large language model), self-supervision, camera-only driving, map-free driving, open-loop evaluation, closed-loop simulation, sensor simulation, generative models, generalization, explanability, multi-task learning.

Autonomous Driving History and Current Status

The speaker discusses the evolution of autonomous driving research, starting from simple neural networks in the 1980s to the end-to-end driving models developed around 2020 by companies like NVIDIA. Early models, demonstrated through videos, exhibited L2-level capabilities, including drifting, which are not suitable for safe passenger transport. The speaker contrasts this with Waymo's L4 system, which operates safely in complex urban environments like San Francisco, navigating pedestrians, cyclists, and traffic signals.

The key to Waymo's success lies in its sophisticated system that encompasses:

  • Perception: Understanding the surrounding environment, including cars, pedestrians, cyclists, traffic lights, and crossroads.
  • Prediction: Forecasting the future states of the world based on current observations.
  • Planning: Determining the optimal driving actions, such as steering, acceleration, and lane changes.

Scaling Challenges and Long-Tail Events

Waymo currently operates in Phoenix, San Francisco, Austin, and Los Angeles, offering rider-only services. The company's ambition is to expand to more cities, including a road trip to 10 cities and even Tokyo, Japan. Scaling presents significant challenges, particularly in handling "long-tail events" – rare and unpredictable scenarios that are difficult to anticipate and program for.

An example of a long-tail event is a traffic control officer overriding a red traffic light, requiring the autonomous system to understand and respond appropriately. Another example is a sudden flock of birds attacking the car. These events, though infrequent, are critical for ensuring safety and require robust generalization capabilities.

Leveraging Foundation Models: Emma and Gemini

The speaker introduces the concept of using foundation models, specifically Google's Gemini, to address the challenges of generalization in autonomous driving. Foundation models are capable of understanding and responding to a wide range of scenarios, even those not explicitly encountered during training.

The speaker presents "Emma," an end-to-end multimodal large language model built on top of Gemini. Emma takes video input from the car's cameras and textual routing information (e.g., "turn left at the next intersection") and outputs future waypoints, indicating the car's desired trajectory.

Key features of Emma:

  • Self-Supervised: Trained using driving logs, where the car's past locations serve as ground truth for future waypoints.
  • Camera-Only: Relies solely on camera input, eliminating the need for LiDAR.
  • Map-Free: Does not require high-definition maps, relying instead on Google Maps for routing.

Performance and Explainability

Emma achieves state-of-the-art performance on the nuScenes open-loop planner benchmark. To improve explanability, the speaker introduces a "channel so reasoning" process, where the model explains its reasoning before outputting the planner. This involves identifying critical objects, predicting their behavior, and determining the appropriate driving meta-decision (e.g., maintain speed, yield, slow down). This approach further improves performance on Waymo's internal open motion dataset.

The speaker also discusses the potential for multi-task learning, where Emma is trained on various tasks simultaneously, including 3D detection, rograph estimation, and visual question answering (VQA). This enhances the model's generalization capabilities and allows it to perform multiple functions within a single framework.

Evaluation and Validation

The speaker emphasizes the importance of evaluation and validation in ensuring the safety and reliability of autonomous driving systems. Different evaluation methods are discussed:

  • Open-Loop Evaluation: Replaying recorded driving videos and assessing the model's performance.
  • Closed-Loop Simulation: Testing the model in a virtual environment where it can interact with the world.
  • Real-World Testing: Deploying the model in actual vehicles and observing its behavior.

The speaker highlights the use of generative models, specifically Google's V2, for sensor simulation. This involves generating realistic videos of driving scenarios, which can then be used to evaluate the performance of Emma under various conditions (e.g., rain, different times of day).

Conclusion

The presentation concludes by emphasizing the potential of foundation models, such as Gemini, to revolutionize autonomous driving. By leveraging these models, Waymo aims to improve the generalization capabilities of its systems, enabling them to handle a wider range of scenarios and scale to new cities. The speaker highlights the ongoing research efforts to adapt and integrate foundation models into Waymo's autonomous driving technology, with the ultimate goal of creating safer and more reliable transportation solutions.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.