Chelsea Finn: Building Robots That Can Do Anything

Y CombinatorAbout 5 min readJul 23, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • General-purpose robots
  • Foundation models for robotics
  • Scale in robot learning
  • Teleoperation
  • Imitation learning
  • Pre-training and fine-tuning
  • Vision Language Models (VLMs)
  • Diffusion models
  • Tokenized actions
  • Diverse data
  • Mobile manipulation
  • Hierarchical vision language action models
  • Synthetic data generation
  • Open-ended prompts and interjections

The Problem: Specialization vs. Generalization in Robotics

The core issue in robotics is the need to build entire companies around specific applications (logistics, lab automation, etc.). This is due to the requirement for custom hardware, software, and movement primitives for each task. This contrasts with the success of foundation models in language, where general models are adapted for specific tasks like coding assistance. The goal is to develop a general-purpose model enabling any robot to perform any task in any environment.

The Role of Scale: Necessary but Not Sufficient

Drawing from the success of language models, the importance of scale is discussed. However, simply scaling up data isn't enough.

  • Industrial Automation Data: Massive scale but lacks diversity for general tasks.
  • YouTube Data: Massive scale but challenges in usability and embodiment gap between humans and robots.
  • Simulation Data: Massive scale but lacks realism and has a reality gap.

Scale is necessary for generalization but subordinate to solving the problem itself.

Physical Intelligence's Approach: Large-Scale Real Robot Data

Physical Intelligence focuses on developing physical intelligence with large-scale real robot data. An example is shown of a teleoperated robot lighting a candle. The goal is to enable robots to perform dextrous, long-horizon tasks, succeed in novel environments, and respond to open-ended prompts.

Case Study: Laundry Folding Robot

A Pi Zero foundation model was trained to unload a dryer and fold laundry. This is presented as a very difficult problem due to the variability in clothing.

Step-by-Step Development:

  1. Start Simple: Fold a single-size, single-brand shirt.
  2. Data Collection: Teleoperation.
  3. Model Training: Imitation learning with a 100 million parameter model mapping images to joint positions at 50 Hz.
  4. Incremental Hardening: Introduce crumpled shirts, then laundry baskets, variable sizes, and shorts.
  5. Breakthrough: Pre-training on all data, then fine-tuning on curated, high-quality demonstration data.

Key Findings:

  • Pre-training and fine-tuning significantly improved performance.
  • The robot could fold five items in a row and stack them.
  • The model was then scaled up using a 3 billion parameter open-source vision language model (Polygeemma).
  • The model takes images and a language command as input and predicts 50 actions into the future using a diffusion head.
  • The pre-trained model was fine-tuned with the same post-training recipe.
  • The robot was able to fold novel clothing items (shorts, V-neck shirts, shirts with buttons).
  • The robot could handle interruptions due to the neural network architecture.

Quantitative Results:

  • Pre-training and post-training significantly outperformed training only on curated data or all data.
  • The same recipe was applied to other tasks (cleaning a table, scooping coffee beans, constructing a cardboard box, lighting a candle).

Takeaways:

  • Independently develop post-training and pre-training.
  • Training on all data doesn't work for complex tasks.
  • Gradually increase the complexity of the task.

Succeeding in Novel Environments: Diverse Data Collection

The limitation of training and testing in the same environment is addressed. The solution is to collect diverse data.

  • Data was collected in homes across San Francisco and diverse mock kitchens and bedrooms (over 100 unique rooms).
  • Mobile manipulation data (tidying bedrooms and kitchens) accounted for only 2.4% of the pre-training mix.
  • The model was also trained on static manipulation data and high-level instructional data.

Challenge:

  • The model initially ignored language instructions.

Solution:

  • Inspired by VLMs, the action head (using diffusion) was modified to prevent deterioration of pre-trained knowledge.
  • Tokenized actions were predicted.
  • The gradient from the randomly initialized diffusion head was stopped.

Results:

  • Faster training.
  • Improved language following (80% follow rate vs. 20%).
  • The model was tested in three Airbnbs it had never been in before.
  • The robot could close cabinets, put away dishes, and clean up spills.

Quantitative Results:

  • Excluding data from static robots reduced performance significantly.
  • Increasing the amount of data from diverse environments increased performance.

Failure Modes:

  • Items not fully in the drawer.
  • Driving over objects.
  • Difficulty picking up thin objects.
  • Confusing the oven for a drawer.

Takeaway:

  • Diverse data enables robots to follow instructions in novel environments.

Responding to Open-Ended Prompts and Interjections: Hierarchical Models and Synthetic Data

The goal is to enable robots to respond to open-ended prompts and interjections, similar to language models.

Approach:

  • Hierarchical vision language action models.
  • A high-level policy breaks down prompts into intermediate verbal responses and atomic language commands.
  • A low-level model executes the commands.

Challenge:

  • Collecting a large number of human-robot interactions is difficult.

Solution:

  • Generate synthetic data using language models to re-label existing robot data with hypothetical human prompts.
  • Train the high-level policy on these synthetic prompts.

Results:

  • The robot could follow prompts like "Make me a ham and cheese sandwich" or "Make me a vegan sandwich, I don't like pickles."
  • The robot could clean up only the trash but not the dishes.
  • The robot could respond to interjections like "Get me something sweet that's not in the basket."

Quantitative Results:

  • The system outperformed existing foundation models in following instructions.

Takeaway:

  • Synthetic data enables robots to respond to open-ended prompts and interjections.

Synthesis/Conclusion

The presentation outlines a path towards developing general-purpose robots by leveraging foundation models, large-scale real-world data, and techniques like pre-training, fine-tuning, and synthetic data generation. The laundry folding robot serves as a compelling example of the progress made. The key takeaways are the importance of diverse data, the benefits of hierarchical models, and the potential for robots to operate in novel environments and respond to complex instructions. While challenges remain, the progress suggests a promising future for physical intelligence.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.