Key Concepts
- Generalization in Robotics: Achieving diverse skills across different environments, tasks, embodiments, and collaboration with humans.
- Passive Data: Image and video datasets used for pre-training robot skills or model parts.
- Artificial Visual Cortex (AVC): A visual perception model pre-trained on out-of-domain video data and fine-tuned for downstream tasks.
- Cortex Bench: A benchmark consisting of 17 tasks across seven benchmarks to evaluate visual perception models for robotics.
- World Models: Action-conditioned models pre-trained on data to predict future states based on actions.
- Habitat Simulator: A visually realistic simulator for training home robots, including rearrangement tasks and social navigation.
- Open EQa: A benchmark for embodied question answering in home environments.
- Partner: A benchmark for multi-agent reasoning and planning, involving a robot and a human agent.
- Digit 360: A vision-based tactile sensor with a silicon fingertip used for dextrous manipulation.
- Dex Gen: A controller trained in simulation to project desired motions into a safe space for the robot hand during teleoperation.
Data-Driven Robotics at Meta FAIR: Generalization and Adaptation
Introduction
The speaker discusses robotics research at Meta FAIR, focusing on generalization and adaptation for robots to perform diverse skills in real-world environments. The presentation covers a deep dive into a specific project and an overview of broader robotics efforts at Meta.
The Challenge of Generalization
Humans possess remarkable generalization capabilities, performing diverse tasks with infinite variations. The key question is how to enable robots to learn these diverse skills in various home environments. The speaker's current focus is on training models that generalize as much as possible before adapting them autonomously after deployment. Generalization in robotics is complex, encompassing different home environments, tasks, embodiments, and human collaboration.
Data, Architectures, and Algorithms
The speaker polls the audience on the most critical factor for achieving generalization: more robot data, better architectures, or better algorithms. The majority believes more robot data is the most important. Meta is heavily investing in data collection, categorized into:
- Passive Data: Image and video datasets.
- Large-Scale Simulation: Reinforcement learning or behavior cloning in simulation with sim-to-real transfer.
- Teleoperation Data: Data collected through human teleoperation of robots.
Deep Dive: Artificial Visual Cortex (AVC)
The speaker details a project focused on utilizing passive data to train diverse robot skills. The approach involves self-supervised pre-training of visual representations on out-of-domain video data, followed by fine-tuning for specific downstream tasks.
Cortex Bench Benchmark
Before developing AVC, the team created Cortex Bench, a benchmark to evaluate existing visual perception models. It consists of 17 tasks across seven benchmarks, covering different observation and action spaces, goal specifications, and simulators (Habitat and MuJoCo). Models like MVP, CLIP, and R3M were evaluated, revealing that no single model outperformed all others.
AVC Training and Results
AVC was trained using a ViT architecture and masked autoencoding (MAE) as a pre-training objective. The team explored different pre-training datasets:
- Ego4D: Used as a base dataset.
- Manipulation Datasets: Something-Something, Epic Kitchens, 100 Days of Hands.
- Navigation Datasets: RealEstate10K, OpenHouse.
- ImageNet: A large image dataset.
Key findings:
- Scaling Model Size: Larger ViT models (ViT-Large) outperformed smaller models (ViT-Base).
- Scaling Datasets: More data was generally helpful.
- Cross-Domain Diversity: Combining manipulation and navigation datasets was more effective than using only manipulation datasets.
- The model trained on Ego4D + manipulation + navigation + ImageNet with ViT-Large performed best on the benchmarks.
- Pre-training on video datasets significantly improved performance compared to randomly initialized models.
Adaptation and Hardware Evaluation
The team then explored adapting AVC through end-to-end fine-tuning and MAE-based adaptation using in-domain data. Fine-tuning generally yielded better results. The approach was also evaluated on hardware with five manipulation tasks:
- TriFinger (CLA helped)
- Random reaching with a Franka arm
- Picking up a bottle with a Franka arm
- Toaster button manipulation
- Opening a drawer
Few-shot imitation learning was used, and AVC outperformed internal baselines.
Future Directions
The speaker is now focusing on action-conditioned world models, pre-training them on data to incorporate action data.
Large-Scale Simulation and Reinforcement Learning
Meta FAIR is also developing visually realistic simulators and benchmarks to train and evaluate high-level reasoning policies, particularly in simulation, with the goal of transferring them to the real world.
Habitat Simulator
Habitat 2.0 (2021) focused on training home robots to rearrange their habitat (pick and place tasks). The simulator, datasets, and assets are open-sourced. Habitat 3.0 focuses on social rearrangement and collaboration with humans.
Mobile Manipulation with Spot Robot
One project involved training a Spot robot to perform general-purpose mobile manipulation tasks (pick and place). The robot was trained in simulation using the Habitat simulator and HM3D/ReplicaCAD datasets. Basic skills (navigation, pick, place) were trained with generalization in mind, and a high-level policy coordinated these skills. The robot was tested in a real-world Fremont apartment.
Recent Benchmarks: Open EQa and Partner
- Open EQa: An agent is dropped into a home environment and asked open-vocabulary questions. The agent must navigate to the location or answer from memory. The benchmark includes over 1600 human-generated questions and an automatic language model-powered evaluation protocol.
- Partner: Enables multi-agent reasoning and planning, involving a robot and a human agent. The tasks involve cleaning the home environment, with the robot adapting to the human's actions. The benchmark comprises 100,000 natural language tasks generated using LLMs.
Teleoperation and Tactile Sensing
Meta FAIR is focused on teleoperation for dextrous manipulation tasks, emphasizing the need for real-world data and tactile feedback.
Tactile Sensing with Digit 360
The Digit 360 is a vision-based tactile sensor with a silicon fingertip. Meta is developing a family of visual tactile sensors and learning general touch representations that can take in tactile images from any of these sensors. The goal is to train an encoder with self-supervised learning for tasks like force estimation, slip detection, and policy learning.
Teleoperation with Dex Gen
The team is working on improving teleoperation with tactile gloves. The approach involves a two-phase process:
- Pre-training in Simulation: Training controllers for multiple tasks in simulation.
- Real-World Adaptation: Projecting the desired motion from a policy or teleoperator into a safe space for the hand using the trained controller (Dex Gen).
The ultimate goal is to combine the Allegro hand with the Digit 360 sensors on a Franka arm mounted on a mobile base to collect diverse data and train better models.
Conclusion
Meta FAIR is pursuing a multi-faceted approach to robotics, focusing on data-driven methods, simulation, and teleoperation. The research aims to achieve generalization and adaptation for robots to perform diverse tasks in real-world environments. The speaker encourages interested individuals to contact them for potential job opportunities.
AI summaries can miss context or contain errors. Check important details against the original video.





