Stanford Robotics Seminar ENGR319 | Autumn 2025 | General Compliant Robot Interaction

Unknown AuthorAbout 7 min readNov 25, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Compliant Robot Interaction: Robots interacting with the environment in a flexible and adaptable manner, especially during contact.
  • Force Torque Sensing: Measuring forces and torques applied to a robot's end-effector or joints.
  • Coin FT: A compact, scalable, and affordable six-axis force torque sensor.
  • Capacitive Force Torque Sensor: A type of sensor that measures force by detecting changes in capacitance.
  • Adaptive Compliance Policy (ACP): A machine learning algorithm that enables robots to learn and adjust their compliance behavior based on force data.
  • UMI (Universal Manipulation Interface): A handheld device for collecting human demonstration data (vision and pose) for robot learning.
  • UMI FT: An enhanced version of UMI equipped with Coin FT sensors for multimodal data collection (vision, depth, pose, and force/torque).
  • Kinesesthetic Teaching: Directly manipulating a robot's end-effector to teach it a task.
  • Force-Informed Control: Robot control strategies that utilize force and torque feedback.
  • Unstructured Environments: Real-world environments that are complex, unpredictable, and contain variations in object properties and clutter.

Coin FT: A Compact and Scalable Force Torque Sensor

This section introduces Coin FT, a novel force torque sensor designed to address the limitations of existing commercial sensors, which are often bulky, expensive, fragile, and difficult to integrate into various robotic platforms, especially those operating in unstructured environments.

  • Problem Statement: Current force torque sensors are expensive (>$1,000, often tens of thousands of dollars), bulky, heavy, and fragile, making them unsuitable for widespread use on smaller robots, robot hands, or in environments with frequent incidental contacts and large impulses.
  • Design Goals: The goal was to create a sensor that is compact, lightweight, slim, affordable, robust, accurate, and tunable for different applications.
  • Coin FT Design:
    • Size: Coin-sized, comparable to a US quarter dollar.
    • Weight: Only 2 grams.
    • Mechanism: Composed of two PCBs with cone-shaped electrodes stacked together, separated by an array of silicon rubber pillars. A shielding layer is placed underneath.
    • Sensing Principle: Leverages the microcontroller's ability to actively reconfigure electrodes to switch between different sensing modes for six axes of force and torque. The core sensing mechanism relies on changes in the relative pose between the two PCBs, governed by the dielectric layer (silicon rubber pillars).
    • Tunability: The mechanical properties of the dielectric layer can be tuned to adjust sensitivity and force range. A compliant layer offers higher sensitivity but lower force capacity, while a stiff layer allows for sensing larger forces but is less sensitive. Parameters like pillar width, material properties, pattern, and number can be adjusted to tune stiffness across different axes.
  • Performance and Robustness:
    • Accuracy: Demonstrates comparable accuracy to expensive ATI sensors (>$10,000) for a given force range, with a material cost of less than $10.
    • Robustness: Highly robust to impacts, including hammer hits, due to its simple design and the shock-absorbing nature of the compliant layer. This makes it suitable for environments with incidental contacts and large impulses.
    • Limitations: Less robust in the tensile direction, with a risk of delamination under very large tensile forces, akin to pulling apart an Oreo cookie.
  • Applications and Impact:
    • Drones: Enables drones to perform contact-based tasks (e.g., collecting samples, window cleaning) in dangerous or inaccessible environments. Its affordability allows for easy replacement after crashes, which are common for drones.
      • Example: A custom drone setup with a Coin FT-equipped end-effector demonstrated an attitude-based force controller to attach a package of sensors to a horizontal surface, adapting its force to successfully complete the task.
    • Wearable Robots/Haptic Devices: Facilitates force-informed control for haptic devices on fingertips, forearms, or arms, leading to more consistent haptic feedback by accounting for individual ergonomic differences.
    • Dexterous Manipulation: Sensorizing multiple fingers of robotic hands (e.g., Allegro hand) to teach robots fine-grained manipulation tasks like plucking a battery from a socket.
      • Example: Studies showed a significant performance jump (80-90% success rates) in six different contact-range tasks when using force-informed policies and compliance control compared to tasks without force feedback.
    • External Adoption: Coin FT is being used at other institutions like Berkeley (Dexter's hands), Switzerland (drones), and UC Santa Cruz (crop manipulation).
  • Future Vision: Open-sourcing Coin FT for researchers and patenting the technology for commercial product development, aiming to make affordable force torque sensors ubiquitous in future robots, especially those interacting in homes.

FT: Learning Compliant Manipulation at Scale from Humans

This section focuses on FT, a framework for enabling robots to learn adaptive compliance behavior and perform contact-based manipulation tasks by learning from human demonstrations, leveraging force torque data.

  • Motivation: While robots are improving, they still struggle with contact-rich tasks in unstructured environments. Human-like adaptive compliance is crucial for safe and effective interaction.
  • Key Requirements for Learning from Demonstration (LfD) with Compliance:
    1. Large-scale Data Collection Platform: Essential for robots to learn from sufficient data.
    2. Powerful Algorithm for LfD and Adaptive Compliance: Capable of learning adaptive compliant behavior from humans.
    3. Compact, Affordable, and Scalable Force Torque Sensor: To support large-scale data collection.
  • Stanford's Solution: UMI FT:
    • UMI (Universal Manipulation Interface): A handheld device that collects vision and pose data from human demonstrations, enabling robots to learn tasks like dishwashing. It's portable, low-cost, and scalable.
    • UMI FT: An enhanced UMI device incorporating Coin FT sensors on each finger, providing finger-level six-axis force torque sensing. It also includes an iPhone for vision, depth, and pose data.
    • Multimodal Data Collection: UMI FT collects vision, depth, pose, and force/torque data simultaneously, enabling robots to learn not just trajectories but also forceful interactions.
  • Adaptive Compliance Policy (ACP):
    • Concept: Compliance control allows robots to behave like springs, enabling safe contact but potentially sacrificing tracking accuracy. ACP addresses this by learning when to be compliant and when to be stiff.
    • Functionality: Based on force data, ACP learns to modulate compliance. For example, when wiping a vase, it can be compliant in the contact direction while remaining stiff laterally to accurately track the contour.
  • Demonstration and Learning with UMI FT:
    • Human Demonstration: A human uses UMI FT to demonstrate a task (e.g., picking up an eraser and wiping a whiteboard). The device provides natural haptic feedback, allowing the human to sense grasp force and contact pressure.
    • Data Capture: UMI FT captures vision, depth, pose, and precise force/torque data from the demonstration.
    • Robot Learning: An enhanced ACP is trained using this multimodal data. The robot learns to replicate the forceful behavior demonstrated, including applying sufficient force for tasks like whiteboard wiping.
  • Generalization and Robustness:
    • Environmental Perturbations: The trained policy demonstrates generalization to out-of-distribution scenarios, such as different table heights, relative board heights, eraser shapes, and drawings, without explicit training on these variations.
    • Baseline Comparisons (Ablation Studies):
      • Force Information, No Compliance Control: Robot fails to adapt quickly to contact force, sometimes ramming surfaces or not pressing hard enough, leading to residue on the whiteboard.
      • No Force Information: Robot fails to grasp objects (e.g., eraser) as grasping is often force-controlled. It overfits to gripper width from training data.
      • Contact Microphone (Tactile Information): Provides good dynamic information but lacks static information. Robots tend to ram surfaces or trigger safety limits due to insufficient static force control.
      • Forceful Interaction Tasks:
        • Skewing Zucchini: Without force information, the zucchini slips off during skewer insertion.
        • Inserting Light Bulb: Without compliance control, the bulb overshoots the slot or slips out of the grasp.
  • Key Takeaway from FT: Combining multimodal data (vision, pose, force/torque) with an adaptive compliance policy is crucial for robots to learn and perform complex, contact-rich manipulation tasks effectively and generalize to variations.

Looking Ahead

The presentation concludes with a vision for future advancements in compliant robot interaction.

  • General Compliant Control with Affordable Sensors: Coin FT will enable a wider range of robot arms, including cheaper ones, to perform compliance control, democratizing this capability beyond expensive industrial robots.
  • Scalable Multimodal Data Collection: Future data collection devices may require different form factors to capture forceful data from robots with multiple fingers, moving beyond the UMI FT concept.
  • Bridging the Data Gap: An ongoing challenge is to bridge the gap between readily available vision/pose data and multimodal force/torque data. Potential solutions include fine-tuning large vision models with tactile information or training residual policies.

The speaker expresses gratitude and opens the floor for questions, highlighting the importance of multiaxial force torque sensing (like Coin FT) for enabling true compliance control, which is not achievable with vision-based force sensing alone.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.