Key Concepts
- Gemini API: Google DeepMind's foundation model API, known for its large context window, multimodality support, and function calling capabilities.
- Agentic Frameworks: Libraries and tools like LangGraph, CrewAI, and Pedantic AI that simplify agent development by providing pre-built components and structures.
- Context Engineering: The process of providing the right information at the right time with the right tools to an LLM to solve a user's task.
- Tool Use/Function Calling: A capability that allows LLMs to interact with external tools and APIs to perform specific tasks.
- Deep Research Agent: An agent designed to perform in-depth research on a given topic, often involving multiple iterations of searching, summarizing, and critiquing.
- Agentic Patterns: Reusable design patterns for building agents, such as tool use, self-reflection, and loop agents.
- Gemini CLI: A command-line interface for interacting with Gemini models, offering features like custom commands, MCP server support, and IDE integration.
- Model Agnostic Prompt: A prompt that can be used with different LLMs without significant loss in performance.
- MCP (Model Code Playground) Server: A server that allows Gemini CLI to execute code in a sandboxed environment.
Gemini as a Model for Agentic Applications
- Unique Capabilities: Gemini stands out due to its 1 million token context window, strong multimodality support (image, audio, video input/output), and reasoning capabilities.
- Extended Reasoning Time: Gemini 2.5 Pro and Flash provide more time for the model to compute and understand complex use cases, beneficial for customer support agents.
- Context Engineering: Gemini's large context window and multimodality support enable developers to provide more comprehensive information to the model, improving its ability to solve user tasks.
- Function Calling: Gemini API supports function calling and structured output (JSON), facilitating interaction between the model and the environment.
- Accessibility: Gemini API is easily accessible through AI Studio, allowing developers to quickly prototype agentic applications with API keys.
Open Models vs. Gemini
- Open Models: Offer benefits in constrained environments where local execution is required.
- Compute Requirements: Open models require more compute resources to run, potentially increasing upfront costs.
- Validation: Start with the easiest path (Gemini API) to validate the proof of concept before investing in GPUs or TPUs for open models.
- Customization: Open models may be considered for higher levels of customization or specific requirements after initial validation with Gemini API.
Smaller Models and Gemma
- Excitement for Small Models: Smaller open models like Gemma 3 270M are promising for edge devices and smart home assistants.
- Trend in Model Sizes: There's a cyclical trend of models becoming smaller and more efficient while maintaining similar capabilities.
- Gemma 3 270M: Requires only 500MB of space, making it suitable for phones, computers, and Raspberry Pi devices.
- Use Cases: Gemma 3 270M is well-suited for information retrieval, structured output generation, and classification tasks.
- Favorite Open Model: Gemma 3 12B is a fast and capable model for local use on a MacBook Pro.
Agentic Frameworks
- Collaboration: Google is collaborating with around 20 startups and open-source libraries to make Gemini more accessible to developers.
- Benefits: Agentic frameworks accelerate prototyping, provide community support, and offer examples for inspiration.
- Considerations: Frameworks address monitoring, observability, and other aspects of production agent development.
- Personal Preference: Choice depends on personal coding style, company standards, and specific use cases.
- Framework Focus:
- LlamaIndex: Document processing and enterprise pipelines.
- CrewAI: Multi-agent interaction and coordination.
- Pedantic AI: Working with structured data.
- No Single Perfect Framework: It's recommended to understand different frameworks and choose the one that best fits the project's needs.
Challenges in Agent Development
- Magical Hello World: Easy initial setup can be misleading, as real-world use cases require careful consideration of guardrails and testing.
- Manual Testing Limitations: Manual testing is insufficient due to the non-deterministic nature of LLMs.
- Evaluation Differences: Evaluation methods for agents differ significantly from traditional software.
- Data Drift: User expectations of AI models shift over time, requiring continuous monitoring and optimization.
- Importance of Observability: Observability, monitoring, and evaluation are crucial for building reliable agentic applications.
Deep Research Agent
- Quick Start Example: A quick start example was built using Gemini and the Google Search integration, utilizing LangGraph for composable agents.
- Community Request: The agent was developed based on community requests and feedback.
- Process Sketching: The development process involved sketching out the steps of researching a topic manually.
- Functions for the Graph: Different functions were created for research, scratchpad management, and self-reflection.
- Self-Reflection Pattern: The agent uses the AI to critique its own output, creating a loop between research, report generation, and feedback.
- Agentic Pattern of Loop Agents: The agent uses the agentic pattern of loop agents in this situation.
Agentic Patterns: Tool Use and Self-Reflection
- Tool Use: Utilizes Google Search to connect with the external environment and gather information.
- Self-Reflection: The LLM critiques its own output based on predefined criteria, determining whether to redo the task or proceed.
- Model Control: The model controls the number of iterations performed during research.
- Guardrails: It's important to implement guardrails to prevent the agent from endlessly researching.
Multi-Agent vs. Single-Agent Systems
- Focus on Achievement: The primary focus should be on achieving the desired outcome, rather than strictly adhering to single-agent or multi-agent classifications.
- Looseness of Definitions: There's often a blurred line between single-agent and multi-agent systems.
- Read-Heavy vs. Write-Heavy: Differentiate between read-heavy and write-heavy use cases to determine the appropriate pattern.
- Parallelization: Multi-agent design patterns are beneficial for read-heavy use cases that can be parallelized.
- Conflict Avoidance: Single-agent or sequential patterns are often easier to implement for write-heavy use cases to avoid conflicts.
- Composability: Agentic use cases and frameworks are highly composable, allowing for flexible mixing and matching of components.
Prompt Engineering and Context Engineering
- Research and Inspiration: Research similar examples and extract patterns from other people's work.
- AI-Assisted Prompting: Use Gemini and AI tools to generate initial prompts and iterate on them.
- Experimentation: Try different approaches and see what works best.
- Context Density: Focus on the density of information in the context rather than specific formatting techniques.
- Context Engineering Definition: Context engineering is more related to the kind of information that you pass in the contest in each turn or each iteration that maybe the agent is going through.
- Instruction Following: Gemini is good at instruction following, so include rules, instructions, and context.
- Emphasis on Instructions: Emphasize instructions at the beginning and end of the prompt.
- Output Format: Determine whether a structured output or function calling is needed.
Memory in Agents
- Potential Extension: Adding memory to the deep research agent could be a valuable extension.
- External Persistent Information: Memory allows for including external persistent information, such as user preferences.
- Personalization: Memory can enable personalized experiences based on user history and preferences.
- Evaluation Challenges: It's difficult to get direct feedback on memory improvements and to differentiate them from other changes.
Evaluation of Deep Research Agent
- Minimal Evaluation: Initial evaluation involved personal testing and team feedback.
- Structured Testing: More structured testing with different search queries and LLM-as-judge pipelines would be beneficial.
- Open Deep Research Benchmark: The open deep research benchmark could be used to evaluate the agent.
- Time-Intensive: Evaluating deep research agents is time-intensive.
- Ship Iterate: The approach is to ship, iterate, and gather feedback from real users.
- Dogfooding: Use the agent personally to gain experience and understanding.
- Risk Assessment: Consider the risk level of the model when determining the appropriate testing approach.
- Simulated User Interaction: Use LLMs to simulate user interactions and expose the agent to different scenarios.
Gemini CLI
- Open Source: Gemini CLI is open source, allowing developers to inspect and extend its functionality.
- Internal and External Use: It's used both internally at Google and externally by third-party developers.
- Documentation Work: Gemini CLI is used internally to update documentation and examples.
- MCP Server Integration: Supports MCP servers for executing code in sandboxed environments.
- Custom Commands: Allows users to create custom commands with predefined prompts and contexts.
- IDE Integration: Offers IDE integration with VS Code for viewing changes.
- Cost Management: Can be used with Gemini API keys from AI Studio or Vertex AI projects.
- Token Usage Reports: Provides reports on token usage to help understand costs.
- Weekly Releases: The Gemini CLI team pushes out weekly releases with new features and improvements.
Model Agnostic Prompts
- Challenge: Scaling agents to production is limited by capacity limits for individual models.
- Goal: Being able to swap from say GPT5 to set seamlessly could help with scaling.
- Approach: Focus on setting up a process to optimize prompts automatically for different models and use cases, rather than writing a single model-agnostic prompt.
- Evaluation: Use evaluation sets to measure performance across different models.
- Automated Optimization: Utilize AI and libraries to automate prompt optimization.
- Ongoing Optimization: Continuously update and optimize prompts due to changing user expectations and data drift.
- Vert.x AI Prompt Optimizer: Google Cloud offers services like Vert.x AI Prompt Optimizer to address this issue.
- Living Artifacts: Agents are living artifacts that require continuous improvement and adaptation.
Conclusion
The conversation with Philip Schmidt provided valuable insights into building production-ready AI agents using Gemini and other tools. Key takeaways include the importance of context engineering, the benefits of agentic frameworks, the challenges of evaluation, and the need for continuous monitoring and optimization. The discussion also highlighted the versatility of Gemini CLI and the potential of smaller open models like Gemma. The emphasis on practical approaches, such as starting with the easiest path for validation and iterating based on real-world feedback, offers actionable guidance for developers looking to build successful agentic applications.
AI summaries can miss context or contain errors. Check important details against the original video.