Key Concepts:
- CAN (Controller Area Network) communication protocol
- Multi-threading and pipelining
- Cycle time and jitter
- Synchronization primitives (mutexes, conditional variables, semaphores)
- Logging overhead
- Priority inversion
1. Introduction: The Gap Between Policy and Actuation
The talk addresses the challenges of translating robot control policies into real-world actions, focusing on the software systems that bridge the gap between the controller and the actuators. The speaker emphasizes that issues that appear to stem from the control policy may actually originate in the underlying software architecture. The goal is to diagnose and resolve these software-related problems to ensure reliable robot performance.
2. Building a Toy Robotics Architecture
- General Architecture: The speaker outlines a basic robot architecture consisting of actuators, a CPU (potentially with a hybrid accelerator), and sensors.
- Communication Protocol (CAN): CAN is chosen as the communication protocol due to its open-source nature, affordability, sufficient data rate, and compatibility with various components.
- Initial Simple Code: The initial code involves receiving data, feeding it to the policy, and sending the output back. The target loop time is 2 milliseconds.
3. Identifying Communication Bottlenecks
- Observed Delay: When the code is deployed on the robot, a gap appears in the loop, extending the cycle time beyond the expected 2 milliseconds.
- CAN Bus Analysis: The speaker analyzes the CAN bus, considering factors such as message size (100 bits per message), the number of messages (10, with 5 sent and 5 received), and the bus speed (1 megabit per second).
- Calculation: With 1000 bits total, the transmission time is calculated to be approximately 0.1 milliseconds per message or 1 millisecond for 10 messages.
- Conclusion: The communication overhead significantly impacts the loop time, explaining the observed 1-millisecond gap.
4. Solution 1: Accepting the Delay
- The simplest solution is to accept the delay, resulting in a 3-millisecond loop time. However, this is not ideal for high-performance systems.
5. Solution 2: Multi-threading and Pipelining
- Goal: To work around the 1-millisecond communication delay and achieve the target 2-millisecond loop time.
- Decomposition: The loop is divided into three components: TX (transmission), Policy (policy execution), and RX (reception).
- Parallelization: Communication (TX and RX) is run in separate threads from the policy execution.
- Staggering: The RX, Policy, and TX tasks are staggered to overlap their execution. The next set of data is received while the current policy is being executed. The data from the last policy is transmitted when the next iteration starts.
6. Addressing Jitter and Desynchronization
- Observed Stuttering: After deploying the multi-threaded system, the robot exhibits stuttering behavior and erratic actuator movements.
- External Transceiver: An external transceiver is connected to the CAN bus to capture raw data.
- Data Analysis: The data is analyzed using tools like
can dumpon a separate host computer (e.g., a laptop). - Cycle Time Plot: A cycle time plot is used to visualize the time since the last message. Ideally, the plot should show a straight line around the 2-millisecond mark.
- Issue Identification: The analysis reveals inconsistent message timing, with some messages arriving late (e.g., 4-millisecond gap) and others arriving almost immediately after the previous one.
7. Root Cause Analysis: TX Side Desynchronization
- Policy Execution Time: Policies may take varying amounts of time to execute, leading to missed transmission deadlines.
- Queuing: When a policy takes longer than expected, the corresponding message is queued.
- Simultaneous Transmission: In the next iteration, both the queued message and the current message are transmitted simultaneously, causing jitter.
8. Addressing TX Side Issues
- Synchronization Primitives: Synchronization primitives (mutexes, conditional variables, semaphores) are recommended to synchronize the TX and RX threads.
- Padding: If synchronization primitives are not available (e.g., in real-time OS or microcontrollers), padding can be added to provide a cushion for desynchronization.
9. Root Cause Analysis: RX Side Desynchronization
- Delayed RX Thread: If the RX thread is delayed, the policy will operate on old data.
- Skipped Data Processing: The policy skips data processing, leading to a "catching up" behavior in the actuators.
10. Addressing RX Side Issues
- Synchronization Primitives: Synchronization primitives are again recommended to ensure proper synchronization of the RX thread.
- Padding: Padding can also be used to mitigate RX side desynchronization.
11. Additional Considerations: Logging Overhead
- Logging to Disk: Excessive logging can block the main control loop, causing the robot to freeze. For example, a Raspberry Pi with an SD card might freeze for 30 milliseconds during logging.
- Solution: Offload logging to a separate CPU to avoid blocking the main control loop.
12. Microcontroller Logging
- UART Logging: On microcontrollers, logging via UART can take a significant amount of time (on the order of milliseconds).
- Packet Drop Issue: Dropping a packet and logging the event can take so long that the next packet is also dropped, leading to a cascade of packet drops and a blackout on the CAN bus.
13. Priority Inversion
- Kernel Interaction: The process of receiving data from the kernel involves multiple steps, including interrupts and kernel processes.
- Priority Boosting: Boosting the priority of user processes too high can block the kernel, preventing it from delivering data.
- Starvation: This priority inversion can cause the system to drop out for extended periods (seconds).
- Solution: Carefully manage process priorities to ensure that the kernel can run and deliver data to user processes.
14. Conclusion
The talk emphasizes the importance of considering the entire software system when designing robotic systems. It covers various techniques for optimizing communication, synchronizing threads, managing logging overhead, and avoiding priority inversion. The speaker highlights the need to understand the interplay between hardware and software to build high-performance robotic systems.
15. Recap of Key Points
- Pipelining: Reduces cycle time by overlapping tasks.
- Synchronization: Prevents jitter caused by desynchronized threads.
- Logging Strategies: Avoid blocking the system during logging.
- Priority Inversion: Prevents starvation by properly managing process priorities.
AI summaries can miss context or contain errors. Check important details against the original video.