How to See What Your AI Agents Are REALLY Doing (ft. Adam Silverman)

Arseny ShatokhinAbout 8 min readJul 12, 2025Watch original
THE SUMMARYAI-generated

AI Agent Observability with Adam Silverman: A Deep Dive

Key Concepts:

  • AI Agent Observability: Monitoring and understanding the performance, behavior, and internal workings of AI agents.
  • Agent Ops: A platform for AI agent observability, providing tools for tracking cost, latency, success/failure rates, and ROI.
  • LLM (Large Language Model): The core AI model powering the agent's decision-making and text generation.
  • Frameworks: Software libraries and tools used to build AI agents (e.g., Agents SDK, CrewAI, Langchain, LlamaIndex, Autogen).
  • Tool Calls: Interactions of the AI agent with external tools, APIs, or services.
  • Latency: The time it takes for an AI agent to respond to a request.
  • ROI (Return on Investment): A metric to quantify the financial benefits of using AI agents.
  • Software Development Life Cycle: The process of planning, creating, testing, and deploying software applications, including AI agents.

Why AI Agent Observability Matters

  • Reliability and Scalability: AI agents are often unreliable initially and require extensive testing to ensure they function correctly and can scale to production environments.
  • Understanding Agent Behavior: Observability provides insights into how agents succeed, fail, and whether they are working towards the intended goals.
  • Beyond Demos: While demo videos on platforms like YouTube may showcase impressive agent capabilities, real-world production deployments require robust monitoring and guardrails.
  • Coding in the Dark: Developing AI agents without observability is akin to coding in the dark, lacking the necessary visibility into their operations.
  • Essential Metrics: Tracking cost, latency, success/failure rates, and tool calls is crucial for understanding and optimizing agent performance.

Use Case: Travel Company and Autogen

  • Custom Solution: A large travel company built its own customer support agent using Autogen from Microsoft instead of using off-the-shelf solutions.
  • Debugging Challenges: The engineering team struggled to debug the agent effectively, lacking auditability and a way to monitor its performance.
  • Agent Ops Implementation: Integrating Agent Ops took less than 5 minutes and provided a comprehensive view of the agent's activities.
  • Productivity Boost: The engineering team reported a 2x increase in productivity due to the improved monitoring and debugging capabilities.
  • Comprehensive Tracking: Agent Ops tracks not only LLM calls but also tool calls, network interactions, and API calls, providing a complete picture of the agent's behavior.

Key Metrics to Track in Production

  • Latency: Crucial for voice and text-based agents to ensure a responsive and enjoyable user experience.
  • Cost: Monitoring LLM and tool usage costs is essential for scaling agents to handle millions of sessions per day. Caching responses can improve latency and reduce costs.
  • Failure Rates: Tracking and being notified of failures (e.g., context window blowouts, API failures) is critical for maintaining agent reliability.
  • ROI Calculator: Quantifying the financial benefits of using AI agents, such as cost savings compared to traditional methods.

Calculating ROI for AI Agents

  • Task-Based Analysis: Comparing the cost of a task performed by a human versus an AI agent (e.g., writing a blog post).
  • Depreciation of Agent Costs: Depreciating the initial cost of building an agent over its lifetime, factoring in LLM costs, API calls, and other expenses.
  • Job Role Automation: Identifying tasks within a job role that can be automated by an AI agent and calculating the cost savings.
  • Sales Optimization: Tracking metrics like cost per lead or customer acquisition cost to determine the impact of AI agents on sales performance.

Best Practices for Implementing Observability

  • Start for Free: Agent Ops offers a free tier to get started quickly.
  • Comprehensive Documentation: Detailed documentation is available to guide users through integrations.
  • Framework Support: Agent Ops supports various frameworks, including Agents SDK, CrewAI, Langchain, LlamaIndex, and custom agents.
  • Unified View: Agent Ops provides a single pane of glass for monitoring data from different frameworks and LLMs.
  • Early Implementation: Integrate observability from day zero to track performance, cost, and latency across different frameworks.

Observability for Small Businesses

  • Productivity Gains: Even solo developers or small teams can benefit from the productivity gains achieved through faster debugging.
  • Cost-Effective Solution: Agent Ops offers an inexpensive solution that can significantly reduce development time.

Agent Ops Differentiation

  • Switzerland Approach: Agent Ops integrates with every LLM provider, agent framework, and tooling, offering a comprehensive solution.
  • Focus on Agent Observability: Agent Ops is specifically tailored for agent observability, tracking tool use and multi-agent tendencies across different LLM providers.
  • Upcoming Features:
    • MCP (Multi-Chain Processing) Observability
    • Chat with Agentic Runs (Beta)
    • Hosting for Agents (Early Access): A hosting platform for agents that integrates with Agent Ops for monitoring.

Common Mistakes in Agent Development

  • Overambitious Projects: Developers often try to automate entire roles or job functions instead of starting with smaller, simpler tasks.
  • Lack of Monitoring: Failing to monitor the agent's performance makes it difficult to track reliability and identify issues.

Recommended Starting Use Cases

  • Chatbots: Easy to get started with and numerous examples available.
  • Niche Automation: Identify small, repetitive tasks within specific industries that can be automated to save time and money.
  • Customer-Driven Development: Talk to potential customers to understand their needs and build agents that address those needs.

Emerging Trends in AI Agent Observability

  • Framework Integration: Continued focus on working with all major agent frameworks.
  • Agent Variations: Supporting the ability to run hundreds of variations of the same agent to optimize performance.
  • Increased Usage: Growing adoption of AI agents across various industries.
  • Tooling Advancements: Improving the experience of seeing and understanding the tools that agents are accessing.

Agent Ops Funding and Focus

  • Debugging as a Core Need: Agent Ops was created to address the challenge of debugging AI agents, recognizing that without debugging, agents cannot be reliably deployed to production.
  • Switzerland Approach: Agent Ops aims to be a neutral platform that works with all LLMs and frameworks, providing a unified view of agent performance.

Competing with LLM Provider Observability Tools

  • Multi-LLM Support: Agent Ops supports multiple LLMs (OpenAI, Anthropic, Gemini, local models), catering to organizations that use a variety of models.
  • Tool Use Tracking: Agent Ops tracks the use of external tools, providing a more comprehensive view of agent behavior than LLM-specific tools.

Lessons Learned Building Agent Ops

  • Risk-Taking: Early on, the company was too conservative and should have taken more risks on promising opportunities.
  • MVP Before Funding: Building a minimum viable product (MVP) that people are willing to pay for is essential before raising funding.

Advice for Entrepreneurs and Founders

  • Deep Problem Understanding: Have a deep understanding of the problem you are solving.
  • Find Like-Minded Investors: Seek investors who share your vision and are willing to take a bet on your company.
  • Build Relationships with Investors: Develop good relationships with investors and seek their advice and support.
  • Choose the Right Partners: Select investors who can be helpful beyond just providing capital.

Location: San Francisco vs. Remote

  • San Francisco Advantages: Access to early adopters, technologically visionary individuals, and a high concentration of engineering talent.
  • Remote Work Considerations: While remote work is possible, being in San Francisco provides unique networking and fundraising opportunities.
  • San Francisco Disadvantages: High living expenses, high taxes, and challenges with visa applications for international talent.

The Evolving Role of Software Engineers

  • AI Code Generation: AI tools are increasingly capable of generating code, but human review and oversight remain essential.
  • Architectural Expertise: The best engineers will focus on architecture, problem-solving, and ensuring the quality of production-grade applications.
  • AI-Assisted Development: Entry-level engineers will increasingly rely on AI tools to assist with coding tasks.

The Future of AI Agents

  • Natural Language Interface: AI agents will become more accessible through natural language interfaces within existing platforms like Slack and Teams.
  • Easier Building Process: Tools like Composio and merge.dev will simplify the process of connecting applications and building agents.
  • Automatic Agent Creation: In the future, AI agents may be automatically built based on user behavior and needs.

Regulatory Compliance in AI

  • Open Source Concerns: Over-regulation of AI could hinder innovation and give other countries an advantage.
  • Privacy and Compliance: Prioritize privacy and compliance, especially when dealing with sensitive information.
  • Focus on Innovation: Regulations should aim to protect consumers while fostering innovation and leadership in the AI space.

Advice for Getting Started in AI

  • Watch YouTube: Utilize YouTube as a valuable resource for learning about AI.

Agency Swarm, Linkfuse, and Agent Ops Integration Tutorial

  • Agency Swarm: A framework for building AI agent swarms.
  • Linkfuse: A platform for evaluating and improving LLM applications.
  • OpenAI Tracing: OpenAI's built-in observability tools.
  • Integration Steps:
    • Install the new version of Agency Swarm using the beta.4 tag.
    • Import necessary libraries.
    • Define API keys.
    • Define tools for the agents.
    • Create agents using the Agency Swarm framework.
    • Set up tracing and observability using OpenAI tracing, Linkfuse decorators, and Agent Ops initialization.
    • Run the agents and observe the traces in the respective dashboards.
  • OpenAI Tracing Dashboard: Provides a clear view of function outputs, user messages, instructions, and tool calls.
  • Agent Ops Dashboard: Shows metrics like success/failure rates and trace duration, and provides a similar view of the agent workflow.
  • Linkfuse Dashboard: Offers additional features like human annotation, LLM as a judge, and custom dashboards.

Conclusion

AI agent observability is crucial for building reliable, scalable, and cost-effective AI agents. Platforms like Agent Ops provide comprehensive tools for monitoring agent performance, tracking key metrics, and debugging issues. By implementing observability from the beginning, developers can gain valuable insights into agent behavior and optimize their performance for real-world applications.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.