Why (Senior) Engineers Struggle to Build AI Agents — Philipp Schmid, Google DeepMind

AI EngineerAbout 4 min readMay 31, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • Agentic Workflow: A paradigm shift from deterministic, step-by-step software execution to goal-oriented, iterative task completion.
  • Semantic Context: Moving away from rigid data structures (Booleans/flags) toward LLM-driven understanding of intent and meaning.
  • Non-deterministic Systems: The reality that agents may take different paths to achieve the same outcome.
  • Evals (Evaluations): A shift from binary unit testing to probabilistic success measurement.
  • Agent-Ready APIs: Designing tools with self-documenting, semantic interfaces that LLMs can interpret without human developer context.

1. The Paradigm Shift: Traffic Controller vs. Dispatcher

Philip (DeepMind) highlights that traditional software engineering is akin to being a traffic controller—defining every road, speed limit, and signal. Building agents, however, is like being a dispatcher: you define the goal (e.g., "get to London"), but the agent determines the specific route (train, plane, or car) based on the context. The focus shifts from rigid workflows to iterative loops of instruction, observation, and adjustment.

2. Five Core Differences in Agent Development

I. Text as the New State

Traditional software relies on structured data (Booleans, enums). Agents operate on semantic meaning.

  • Example: A deep research agent can accept a plan and simultaneously receive nuanced, natural language constraints (e.g., "focus on the US market, ignore California") without needing a hard-coded schema for every possible user preference.

II. Handing Over Control

Developers must move away from rigid, stateful workflows (e.g., a fixed customer support "cancel subscription" flow).

  • The Challenge: Modeling every possible user intent is impossible.
  • The Solution: Trust the LLM to handle dynamic interactions. If a user changes their mind during a cancellation flow, the agent should be capable of pivoting its intent rather than being trapped in a predefined, deterministic path.

III. Errors as Inputs

In traditional systems, errors often trigger a restart. In long-running agentic processes (which may take 5–15 minutes), restarting is computationally expensive and loses context.

  • Methodology: Treat errors as "normal inputs." Feed the error back into the model so it can self-correct or find a workaround, maintaining the flow rather than failing the entire process.

IV. From Unit Tests to Evals

Because agents are non-deterministic, the same input does not guarantee the same output.

  • The Shift: Move from binary unit tests (Pass/Fail) to Evals.
  • Measurement: Use "LLM-as-a-judge" or human experts to qualify success. Since agents might take different paths to reach the same goal, success should be measured by the outcome rather than the specific steps taken.

V. Agents Evolve, APIs Don't

Standard APIs are built for human developers who possess years of context. Agents, however, only see function schemas and docstrings.

  • Requirement: APIs must be "agent-ready." This means using highly descriptive, self-documenting semantic interfaces. A method named delete_item is insufficient; the tool definition must explicitly explain the "why" and "how" for an LLM to use it effectively.

3. Key Arguments and Perspectives

  • Stop Fighting the Model: Developers often try to force LLMs into rigid, step-by-step workflows. Philip argues this is counterproductive; instead, define the goal and allow the model to navigate the path.
  • Design for Recovery: Since agents are not perfect, systems must be built to handle "weird" behavior and recover gracefully.
  • Software is Disposable: The "bitter lesson" is that agentic code will be rebuilt frequently as models improve. Don't over-engineer; build to be replaced.

4. Synthesis and Conclusion

Building agents requires a fundamental shift in mindset:

  1. Trust but Verify: Give the agent autonomy, but implement robust evaluation frameworks to ensure reliability.
  2. Preserve Meaning: Leverage the LLM's ability to process context rather than forcing data into rigid structures.
  3. Iterative Design: Accept that agents are non-deterministic and design systems that treat errors as part of the learning/execution loop.
  4. Semantic Tooling: Ensure all tools and APIs are explicitly documented for machine consumption, not just human developer intuition.

Final Takeaway: The transition from traditional software to agentic systems is a move from deterministic control to probabilistic goal-seeking. Success lies in designing for flexibility, recovery, and semantic clarity.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.