How to Build Reliable AI Agents in 2025

Dave EbbelaarAbout 5 min readJul 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • AI Agents: Automated systems leveraging large language models (LLMs) to perform tasks.
  • LLMs (Large Language Models): AI models trained on vast amounts of text data, capable of generating human-like text, reasoning, and understanding.
  • Deterministic Software: Software that produces the same output for a given input, predictable and reliable.
  • Context Engineering: The process of crafting and managing the information provided to an LLM to ensure accurate and relevant responses.
  • Workflows/DAGs (Directed Acyclic Graphs): A series of interconnected steps or tasks, often involving both code and LLM calls, to achieve a specific goal.
  • Intelligence Layer: The component responsible for making the actual API call to the LLM.
  • Memory: The ability to retain and utilize past interactions to inform current responses.
  • Tools: External functions or APIs that the LLM can call to perform actions beyond text generation.
  • Validation: Ensuring that the LLM's output conforms to a predefined structure or schema.
  • Control: Using deterministic code (if/else statements, routing logic) to guide the flow of the application.
  • Recovery: Implementing error handling and fallback mechanisms to ensure resilience.
  • Feedback: Incorporating human oversight and approval steps to ensure quality and safety.
  • Structured Output: LLM responses formatted in a predefined structure, such as JSON, for easier processing.

1. The Problem: AI Overload and Misdirection

  • The AI space is oversaturated with information, tools, and frameworks, leading to developer confusion and anxiety.
  • Many tutorials and resources are messy, contradictory, and focus on complex agent frameworks rather than foundational principles.
  • The hype surrounding AI agents often overshadows the importance of traditional software engineering practices.
  • The core issue is that many developers are distracted by the "noise" and don't focus on the fundamental building blocks.
  • The speaker aims to provide clarity and help developers focus on the core principles for building reliable AI agents.

2. The Solution: Foundational Building Blocks and Software Engineering Principles

  • The key to building effective AI agents is to understand and utilize seven foundational building blocks.
  • The most effective AI agents are primarily deterministic software with strategic LLM calls.
  • Developers should prioritize solving problems with code and only use LLMs when necessary.
  • LLM API calls are expensive and potentially unreliable, so they should be minimized.
  • Context engineering is crucial for getting accurate and reliable responses from LLMs.
  • AI agents are essentially workflows or DAGs, with most steps being regular code, not LLM calls.

3. Seven Foundational Building Blocks

3.1. Intelligence Layer

  • This is the core AI component where the API call to the LLM is made.
  • It involves sending a user input to the LLM and receiving a response.
  • Example: Using the OpenAI Python SDK to connect to the API, select a model, and send a prompt.
  • The speaker emphasizes that the LLM call itself is straightforward; the complexity lies in the surrounding components.

3.2. Memory

  • LLMs are stateless, meaning they don't remember previous interactions.
  • Memory ensures context persistence by storing and passing conversation history to each interaction.
  • Example: Storing a conversation history as a sequence of messages (user and assistant) and passing it to the LLM with each new input.
  • In real-world applications, conversation history would be stored and retrieved from a database.

3.3. Tools

  • Tools enable LLMs to interact with external systems, such as APIs, databases, and files.
  • The LLM decides whether to use a tool based on the input and context.
  • If a tool is selected, the code executes the tool and passes the result back to the LLM.
  • Example: Providing the LLM with a "get weather" tool that can retrieve weather information for a given city.
  • Tool calling is directly supported by major model providers, eliminating the need for external frameworks.

3.4. Validation

  • Validation ensures that the LLM's output conforms to a predefined structure or schema.
  • LLMs are probabilistic and can produce inconsistent outputs, so validation is crucial for quality assurance.
  • If validation fails, the output is sent back to the LLM for correction.
  • Example: Using Pydantic to define a JSON schema for task information (task, due date, priority) and validating the LLM's output against that schema.
  • Structured output is essential for building reliable systems around LLMs.

3.5. Control

  • Control involves using deterministic code (if/else statements, routing logic) to guide the flow of the application.
  • This prevents the LLM from making every decision and allows for more predictable and reliable behavior.
  • Example: Using an LLM to classify the intent of a user message (question, request, complaint) and then routing the message to the appropriate function based on the intent.
  • The speaker prefers using structured output and intent classification over tool calls for complex systems, as it provides better debugging capabilities.

3.6. Recovery

  • Recovery involves implementing error handling and fallback mechanisms to ensure resilience.
  • APIs can go down, LLMs can return nonsense, and rate limits can be hit, so error handling is essential.
  • Example: Using try/except blocks to catch errors and retry requests with backoff or provide fallback responses.
  • This is standard error handling that would be implemented in any production system.

3.7. Feedback

  • Feedback involves incorporating human oversight and approval steps to ensure quality and safety.
  • Some processes are too complex or sensitive to be fully automated, so human review is necessary.
  • Example: Requiring a human to review and approve an LLM-generated email before it is sent to a customer.
  • This is a basic approval workflow with a human in the loop.

4. Orchestration and Workflow Design

  • The speaker emphasizes the importance of orchestrating these building blocks to create complete workflows.
  • The key is to break down a large problem into smaller problems and solve each sub-problem using the appropriate building blocks.
  • LLM API calls should only be used when absolutely necessary.
  • The speaker recommends checking out his workflow orchestration course for more information on how to combine these building blocks.

5. Conclusion

  • By understanding and utilizing these seven foundational building blocks, developers can build reliable and effective AI agents.
  • The key is to prioritize software engineering principles and minimize reliance on LLM calls.
  • Context engineering, validation, and control are crucial for ensuring quality and reliability.
  • Human oversight and feedback are essential for complex or sensitive processes.
  • The speaker encourages developers to focus on these core principles and ignore the hype surrounding complex agent frameworks.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.