3 ingredients for building reliable enterprise agents - Harrison Chase, LangChain/LangGraph

AI EngineerAbout 4 min readJul 24, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Reliable Agents in the Enterprise
  • Value of Agents (when right)
  • Cost of Agents (when wrong)
  • Probability of Success
  • Deterministic vs. Non-Deterministic Agents
  • Workflows vs. Agents (and the spectrum between them)
  • Observability and Eval (LangSmith)
  • Human-in-the-Loop (HITL)
  • Ambient Agents
  • Sync to Async Agents
  • Agent Inbox

Building Reliable Agents in the Enterprise

The core idea is to explore how to build reliable agents for enterprise use, focusing on factors that contribute to their success and adoption. The discussion stems from observations of developers building agents within enterprises and solution providers aiming to sell agent-based solutions.

The Vision of the Future

The vision is an enterprise populated by numerous agents handling diverse tasks, with humans acting as managers or supervisors coordinating their activities. The key question is how to achieve this vision and which aspects will materialize first.

First Principles: Value, Cost, and Probability of Success

The success of an agent in the enterprise hinges on three fundamental factors:

  1. Value when Right: The greater the value an agent provides when it functions correctly, the higher the likelihood of its adoption.
  2. Cost when Wrong: Conversely, significant costs associated with agent errors reduce the likelihood of adoption.
  3. Probability of Success: The more reliable the agent is, the more likely it is to be adopted.

This can be represented by the equation: (Probability of Success * Value when Right) - Cost when Wrong > Cost of Running the Agent.

Increasing the Value of Agents

  • Focus on High-Value Verticals: Target problems in areas where the value of correct solutions is substantial. Examples include:
    • Legal: Harvey, an agent in the legal space.
    • Finance: Research and summarization tasks.
  • Shift UI/UX for Long-Term, Substantial Work: Move away from quick, simple question-answering towards deep research and extended operations.
    • Deep Research: Agents that run for extended periods.
    • Ambient Agents for Code: Agents that operate in the background for hours.
    • Example: The trend from quick question answering to deep research, and from inline autocomplete to ambient coding agents.

Increasing the Probability of Success

  • Reliability through Determinism: Make agents more deterministic to increase predictability and control, especially in enterprise settings where specific workflows are required.
    • Workflows vs. Agents: Adopt a hybrid approach, combining workflows (deterministic steps) with agentic components (LLM-driven actions).
    • LangGraph: An agent framework that supports this spectrum of workflows and agents.
  • Reduce Perceived Error Bars: Address uncertainty and fear surrounding agent performance by demonstrating reliability and transparency.
    • Observability and Eval (LangSmith): Use tools like LangSmith to visualize agent behavior, benchmark performance, and communicate insights to stakeholders.
    • Example: A user successfully presented LangSmith data to a review panel, reducing perceived risk and expediting approval.

Reducing the Cost When Wrong

  • UI/UX Tricks: Implement design elements that mitigate the perceived risk of agent errors.
    • Easy Reversal: Enable easy reversal of agent actions.
      • Example: Replit's coding agent creates commits for every change, allowing easy reversion.
    • Human-in-the-Loop (HITL): Incorporate human review and approval steps.
      • Example: Opening pull requests (PRs) for code changes allows human oversight before merging.
  • Concrete Examples:
    • Deep Research: Calibrating research goals with follow-up questions and delivering a report for human action, not automated publication.
    • Claude Code: Clarifying questions and code changes on a separate branch with a PR.

Scaling Agents: Ambient Agents

  • Ambient Agents: Agents triggered by events, operating in the background without direct human initiation.
    • Scalability: Ambient agents enable a one-to-many relationship, scaling up the positive expected value.
    • Concurrency: Move from limited concurrent chat sessions to potentially unlimited background processes.
    • Latency: Relaxed latency requirements allow for more complex operations.
  • Ambient Does Not Mean Fully Autonomous: Maintain human oversight through various interaction patterns.
    • Approve/Reject: Explicitly approve or reject tool calls.
    • Edit Tool Calls: Correct agent errors in tool calls.
    • Question Answering: Allow agents to ask for clarification.
    • Time Travel (Human-on-the-Loop): Revert to a previous step and resume with corrections.
  • Sync to Async Agents: An intermediary state where humans initiate tasks but agents operate asynchronously.
    • Example: Factory's "async coding agents."
  • Agent Inbox: A UX pattern for surfacing agent actions requiring approval.
  • Email Example: An ambient agent that listens to incoming emails, processes them, and requires human approval before sending replies or calendar invites.
    • Open Source Example: A personal email agent built and shared on GitHub.

Conclusion

Building reliable agents in the enterprise requires a focus on maximizing value, minimizing cost, and increasing the probability of success. This involves strategic problem selection, deterministic design, transparent operation, and human-in-the-loop mechanisms. The future points towards ambient agents that operate autonomously in the background, triggered by events, but still subject to human oversight and control.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.