Human seeded Evals — Samuel Colvin, Pydantic

AI EngineerAbout 4 min readJul 25, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Agents: Entities with an environment, tools, a system prompt, and a loop for interacting with the environment.
  • Type Safety: Using static typing to ensure code correctness and facilitate refactoring.
  • Pydantic AI: A framework for building AI applications with a focus on type safety and structured data extraction.
  • Validation Errors: Using validation errors from Pydantic models to guide the LLM in correcting its output.
  • Tools: Functions that agents can use to interact with the environment.
  • Dependencies (Deps): Type-safe dependencies for tools, ensuring that tools receive the correct data types.
  • Logfire: An observability platform for tracing and debugging AI applications.
  • MCTP (Mentioned but not covered): Monte Carlo Tree Search.
  • Eval (Mentioned but not covered): Evaluation of AI models.

1. Introduction and the Need for Reliable AI Applications

  • The speaker introduces the topic of building AI applications using Pydantic, emphasizing the importance of reliability and scalability.
  • Despite rapid changes in the AI landscape, the fundamental need for robust applications remains.
  • Building reliable applications is arguably harder with GenAI, whether using GenAI to build the application or using GenAI within the application.
  • Type safety is presented as a crucial aspect of building reliable AI applications, aiding in both bug prevention and code refactoring.
  • Type safety allows coding agents like Cursor to "mark its own homework" by using type checking to validate its actions.

2. Defining Agents and the Agantic Loop

  • The speaker presents a widely accepted definition of an agent, referencing Barry Zang's talk at AI Engineer in New York.
  • An agent interacts with an environment using tools, guided by a system prompt, in a continuous loop.
  • The loop involves calling the LLM, receiving actions, running tools, updating the state, and calling the LLM again.
  • A key challenge is determining when to exit the loop. Strategies include ending when the LLM returns plain text, using "final result tools," or leveraging structured output types from models like OpenAI or Google.

3. Minimal Example of Pydantic AI: Data Extraction

  • A simple Pydantic model with three fields is used to demonstrate structured data extraction from unstructured data.
  • Pydantic AI extracts structured data that fits the person schema from unstructured data.
  • The example showcases how Pydantic AI can extract data from large documents and complex nested schemas.
  • The initial example is a one-shot process: one call to the LLM returns structured data, which is then validated by Pydantic.

4. Leveraging Validation Errors in the Agantic Loop

  • The example is modified to include a field validator that requires the date of birth to be before 1900.
  • When the model returns an invalid date of birth (e.g., 1987), Pydantic AI returns the validation error to the model.
  • The model uses the validation error to correct its output and try again.
  • This technique is effective for fixing simple use cases where models fail validation.
  • In the example, the model makes two calls to Gemini, with the validation error guiding the second call to success.

5. Type Safety and Generic Types

  • The speaker emphasizes the value of static typing in Pydantic AI.
  • The agent function's output type is generic (e.g., Person), ensuring that the returned data is an instance of the specified type.
  • Accessing fields on the output is type-safe, with typing errors occurring if incorrect fields are accessed.
  • The agent function has a second generic type (Deps) for type-safe dependencies to tools.

6. Tools and Type-Safe Dependencies

  • An example is presented using tools for long-term memory: record_memory and retrieve_memory.
  • Tools are registered with the agent using a decorator.
  • The Deps type is used to define dependencies for the tools.
  • The run_context parameter in the tool decorator is parameterized with the Deps type, ensuring type safety when accessing dependencies within the tool.
  • The speaker claims that Pydantic AI is the only agent framework that works this hard to be type-safe.

7. Observability with Logfire

  • The example is instrumented with Logfire, an observability platform.
  • Logfire provides tracing information, showing how long each call took and the pricing information.
  • Logfire helps in debugging and understanding the flow of execution.
  • A demonstration of Logfire's value is shown when the example fails due to a substring mismatch in the memory retrieval tool.
  • Logfire reveals that the retrieve_memory tool failed because the query "your name" was not a substring of the previously recorded memory "user's name is Samuel."

8. Conclusion

  • The speaker concludes by thanking the audience and highlighting the benefits of Pydantic AI for building reliable and type-safe AI applications.
  • The talk emphasizes the importance of type safety, validation errors, and observability in the development process.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.