Building Applications with AI Agents — Michael Albada, Microsoft

AI EngineerAbout 5 min readJul 25, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

AI Agents, Agentic Development, Foundation Models, Tool Use, Orchestration Patterns, Multi-Agent Systems, Evaluation, Observability, Common Pitfalls, Safety, Productivity.

1. Introduction to AI Agents and Agentic Development

  • Main Topic: Building applications with AI agents.
  • Key Points:
    • Michael Alba, Principal Applied Scientist at Microsoft, presents on building applications with AI agents.
    • He has contributed to Security Copilot and Security Copilot Agents.
    • The talk is based on a forthcoming O'Reilly book, with the first seven chapters available on early release.
    • The presentation focuses on slides, but full code examples are available.
  • Overview: The talk covers the promise and obstacles of agentic development, core components for effective AI systems, and common pitfalls.

2. The Promise and Obstacles of Agentic Development

  • Main Topic: The increasing interest and challenges in agentic development.
  • Key Points:
    • A 254% increase in companies describing themselves as agentic in Y Combinator over the last three years indicates growing interest.
    • Agentic benchmarks from academia show that these are hard tasks requiring multiple tool calls and complex environments.
    • While initial prototypes can achieve around 70% accuracy, reaching higher accuracy in complex scenarios is challenging.
  • Data: 254% increase in agentic companies in Y Combinator.
  • Perspective: While the field is promising, perfection should not be expected, and achieving high accuracy is difficult.

3. Defining AI Agents and Their Effectiveness

  • Main Topic: Defining AI agents and emphasizing effectiveness.
  • Key Points:
    • An AI agent is defined as an entity that can reason, act, communicate, and adapt to solve tasks.
    • The foundation model is the base, and additional components enhance performance.
    • Agency is a spectrum, not a binary distinction, and effectiveness is a crucial second axis.
    • Agency should be a tool to solve problems, not a goal in itself.
  • Example: Robotic Process Automation (RPA) has low agency but high efficacy.
  • Argument: Adding agency should maintain a high level of effectiveness; avoid compromising effectiveness for the sake of agency.
  • Perspective: Avoid building agentic systems that are low in efficacy.

4. Tool Use in AI Agents

  • Main Topic: The power and risks of tool use in AI agents.
  • Key Points:
    • Foundation models can output function calls, allowing agents to invoke tools via APIs.
    • Exposing functionality requires discernment and responsibility.
    • The process involves parsing outputted text, invoking tools, receiving observations, and generating final output.
  • Fallacy: Avoid a one-to-one mapping between APIs and tools.
  • Recommendation: Reduce the number of tools exposed at a single time, group them logically, and ensure specific names and descriptions.
  • Empirical Observation: Exposing more tools to an individual prompt decreases overall accuracy due to semantic collision.

5. Orchestration Patterns for Tool Invocation

  • Main Topic: Orchestrating tool invocation in AI agents.
  • Key Points:
    • Keep orchestration simple and use standard workflow patterns.
    • Single chains are preferred for ease of measurement, cost efficiency, and reliability.
    • Branching logic can be applied, with the LLM choosing the path.
    • Full agentic patterns give more power to the model but are harder to measure.
  • Recommendation: If chains become too complex, consider moving to a more agentic pattern or fine-tuning the model.
  • Pattern: Expose tools to update states and apply validation to maintain deterministic business logic.
  • Example: Cyber security incident severity assessment using branching logic.

6. Multi-Agent Systems

  • Main Topic: Scaling with multi-agent systems.
  • Key Points:
    • Break down a single agent system into multiple agents to avoid overwhelming a single prompt with too many tools.
    • Group tools into semantically similar groups and register them with individual agents.
    • Use a coordinator to route tasks to the appropriate agent.
  • Future Direction: Agent-to-agent protocol aims for coordination between agents built by different teams.
  • Challenge: Agent-to-agent protocol is in early stages with technical and security questions to address.

7. Evaluation of AI Agents

  • Main Topic: The importance of rigorous evaluation in AI agent development.
  • Key Points:
    • Invest more in evaluation due to the numerous hyperparameters to choose (number of agents, tools, model, memory).
    • Move towards test-driven development, defining agents in terms of expected inputs and outputs.
    • AI architects and engineers should take ownership of defining what the agent should do.
  • Process:
    1. Take user inputs and run them through the agent.
    2. Human review of outputs.
    3. Add new additions to the evaluation set.
    4. Run the evaluation set through the agent.
    5. Analyze failures, cluster and summarize outputs, and suggest improvements.
  • Tools: Intel Agent (synthetic inputs), Microsoft Pirate (red teaming), Label Studio (evaluation sets), trace, textrad, ds pi (hyperparameter optimization).

8. Observability in AI Agents

  • Main Topic: The importance of observability for understanding AI agent behavior at scale.
  • Key Points:
    • Generative models make it challenging to evaluate systems at scale.
    • Use tools like OpenTelemetry integrations for detailed logs and tracing.
    • Implement clustering and automated summarization to understand failure modes.

9. Common Pitfalls in AI Agent Development

  • Main Topic: Common mistakes and challenges in building AI agents.
  • Key Points:
    • Insufficient evaluation is the biggest limitation.
    • Inadequate or overlapping tool descriptions.
    • Excessive complexity.
    • Lack of a tight learning loop.
    • Agentic systems are a new class of potential vulnerability.
  • Recommendation: Design for safety at every layer, build trip wires and detectors, and fall back to human review in critical cases.

10. Conclusion

  • Main Takeaway: AI agents have the potential to significantly increase productivity.
  • Quote: "Productivity isn't everything, but in the long run, it is almost everything" - Paul Krugman.
  • Perspective: AI agents represent a new design pattern that can help individuals accomplish more.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.