Agents are Robots Too: What Self-Driving Taught Me About Building Agents — Jesse Hu, Abundant

AI EngineerAbout 8 min readNov 24, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • 1% vs. 99% Problem: The realization that in real-world applications, the core model is only a small part of the overall effort, with the majority dedicated to infrastructure, data, and deployment.
  • Embodiment: The concept of a "body" for an agent, which in robotics is physical, and in digital agents includes tools, APIs, terminals, browsers, and operating systems.
  • Offline Stack: The crucial infrastructure for simulation, training, data collection, and human feedback loops that supports agent development and deployment.
  • Open Loop vs. Closed Loop: The distinction between taking an action without feedback and taking an action with real-time observation and recalibration based on the outcome.
  • Implicit Discretization in Time: The design choice of how frequently an agent samples the world and interacts, contrasting the turn-based nature of conversations with the continuous sampling in robotics.
  • Input and Action Spaces: The range of modalities and methods an agent can use to perceive the world (inputs) and to interact with it (actions), including traditional tools, character-level terminal interaction, and even frame-by-frame GUI interaction.
  • Stateless vs. Stateful Processes: The shift from agents that operate in isolated sessions to agents that maintain persistent state, including running processes and file systems.
  • Dagger and Out-of-Distribution (OOD) Problem: The challenge in imitation learning where deviations from training data lead to significant performance degradation, a problem observed in both robotics and agents.
  • Actions Have Consequences: The fundamental paradigm shift from predictive models to models that act and must deal with the repercussions of those actions in a complex environment.
  • Markov Decision Process (MDP): A mathematical framework for modeling decision-making in situations where outcomes are partly random and partly under the control of a decision-maker, comprising states, actions, and rewards.
  • Simulation: The use of simulated environments to represent real-world complexities, allowing for training, evaluation, and playing out counterfactual scenarios.
  • Hill Climbing: An iterative process of improving a complex system by making incremental changes and evaluating their impact on a target metric, often involving guess-and-check.
  • Agentics: The proposed term for the dedicated scientific practice of developing agentic systems, drawing parallels to the established field of robotics.

Parallels Between Robotics and Digital Agents

The talk draws significant parallels between the development of self-driving cars and robotics and the emerging field of digital agents. A core observation is the "1% versus 99% problem." While the AI model itself might represent only 1% of the effort, the remaining 99% is dedicated to the surrounding infrastructure, data pipelines, and deployment mechanisms.

Embodiment and the "Body" of an Agent

  • Robotics: Embodiment is literal, with a "brain" (the AI model) connected to a physical body comprising hardware, sensors, and actuators.
  • Digital Agents: Embodiment is conceptualized as a "digital robot" that includes a suite of tools. This extends beyond simple APIs and MCPs to more advanced interfaces like terminals, browsers, and virtual machines (VMs). This digital embodiment can encompass entire operating systems and persistent file systems, akin to the "hands, arms, and legs" of a physical robot.

The Importance of the Offline Stack

Both fields rely heavily on an "offline stack" for development and continuous improvement. This includes:

  • Simulation environments for training and testing.
  • Data collection and processing.
  • Continuous retraining and monitoring.
  • Human feedback loops.
  • Tooling to support agent development.

The speaker emphasizes that in self-driving, the winning teams were not just those with the best models but those with the most robust offline stacks, enabling faster and more reliable development.

Open Loop vs. Closed Loop Systems

This concept, crucial in robotics, is being applied to agents:

  • Robotics: A closed-loop system involves taking an action (e.g., turning a wheel) and then receiving feedback on the actual outcome (e.g., measuring the car's turn angle) to recalibrate and ensure accuracy.
  • Digital Agents: An open-loop system might involve running a bash command without real-time observation of its completion or output. A closed-loop approach would involve monitoring the process, detecting completion, and potentially exiting early if needed.

Implicit Discretization in Time and Action/Input Spaces

  • Robotics: Explicit design choices are made regarding how to discretize time and the input space (e.g., vision, LiDAR, radar) and action space (e.g., XY coordinates, acceleration, velocity). Sampling rates (e.g., 50 Hz) determine how frequently the agent updates its state and reacts to the environment.
  • Digital Agents: Discretization is often implicit, particularly in conversational agents. The agent waits for a turn, executes a tool, and waits for the entire response. This turn-based interaction, while easy to reason about, limits real-time responsiveness to events like pop-ups or long-running processes.
  • Input/Action Space Nuances: The speaker highlights the Terminus agent from Terminal Bench as an example of advanced interaction, using T-X streams for character-by-character input/output, enabling fine-grained control. This contrasts with traditional agent designs that might only consider tool calls. The analogy is drawn to interacting with a computer at a frame-by-frame level (20 FPS) with mouse clicks and keyboard inputs, as seen in the "Dreamer" paper. The core question is what trade-offs are made with these design decisions.

From Stateless to Stateful Processes

  • Stateless: Agents that operate in isolated sessions, where the state is not preserved after termination (like a video game session).
  • Stateful: Agents that maintain persistent state, similar to a real car that has mass and occupies space. In the digital realm, this translates to agents operating within VMs that have running processes and persistent file stores. This necessitates considering the entire environment, including ongoing Slack messages and the overall state of the world. This has implications for evaluation and simulation.

The Dagger and Out-of-Distribution Problem

This is a well-known issue in both robotics and agents:

  • Imitation Learning: Training models using human demonstrations (similar to SFT) can lead to significant performance degradation when the agent encounters situations slightly outside the training distribution (OOD).
  • Example: A browser agent encountering a pop-up it has never seen during training can become confused and fail. This is a cascading issue where actions have consequences, and the agent must learn to recover from mistakes.

Actions Have Consequences and the Role of Simulation

  • Paradigm Shift: The move from purely predictive models to models that predict, act, and then deal with the consequences of those actions is a fundamental challenge.
  • Complexity of the Real World: The messiness and complexity of the real world necessitate simulation. Simulation allows for representing these complexities and exploring not just single paths but all possible counterfactuals.
  • MDP Framework: The Markov Decision Process (MDP) provides a formal way to conceptualize the agent loop, involving states, rewards, and actions within an environment. This framework is crucial for describing and communicating agent behavior.

The Deceptive Trickiness of Action Models

  • Self-Driving Analogy: Early self-driving efforts focused heavily on perception models (e.g., drawing bounding boxes), assuming that driving would be easy once the world was perceived. This proved to be an oversimplification, with significant hidden complexity in creating action models.
  • Language Models: Similarly, while LLMs can understand and reason extensively, implementing complex plans and tool calls in the real world often leads to failures. Agents may fail to progress, correct mistakes, or handle tool call errors. This highlights that the bulk of the work lies in the action execution and recovery loop.

The Advantage of Predefined Interfaces

The speaker identifies a key reason for the success of self-driving in limited cases:

  • Predefined Human Interface: Cars have well-defined human controls and electronic interfaces, along with built-in telemetry. This provides a structured way to take actions and collect data, making ML development more convenient.
  • Contrast with Other Domains: This is easier than working with less codifiable tasks that require full desktop interaction. When exploring new domains, the presence of such predefined interfaces is a significant advantage.

The Hill Climbing Process

  • Iterative Improvement: This is an iterative process of building and refining complex systems like LMs or agents, where forward progress is not always guaranteed.
  • Guess and Check: In the absence of direct guarantees, progress is made through a process of guessing, experimenting, and hoping to improve a nebulous metric.
  • Self-Driving Approach: A more sophisticated approach involves learning, then simulating, and then deploying. Real-world logs from deployment feed back into the simulation engine, grounding it and providing crucial insights.
  • Importance of Logs: Logs are more valuable than raw metrics. Breaking down failures by category, city, or specific error types provides deeper insights for improvement. This is a core focus of tooling and processes developed for customers.

Current State and Future Outlook

  • Remote Labor Benchmark: The speaker uses a metric from the remote labor benchmark to illustrate that while impressive demos and predictive models exist, end-to-end work completion is still a challenge, comparable to early stages of self-driving.
  • Recap of Key Learnings: The talk summarizes the parallels, including closed-loop systems, discretization, input/action spaces, statefulness, action models, simulation, and the importance of infrastructure.
  • "Agentics": The speaker proposes the term "agentics" to elevate agent development into a dedicated scientific practice, drawing parallels to the coolness and rigor of robotics.
  • Further Reading: Recommendations include concepts like open-loop/closed-loop control, MDPs, fully/partially observable environments, dagger, offline RL, and introductory reinforcement learning texts. The convergence of robotics and agent fields suggests that robotics literature is also highly relevant.
  • Conclusion: Agents are essentially robots that act in the real world, make mistakes, and must recover. These details are critical.

The speaker, Jesse Abundant, can be reached at [email protected] for thoughts or feedback.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.