How to build a Data Science agent with ADK

Google for DevelopersAbout 5 min readApr 23, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • AI Agents: Plan, reason, and execute tasks on your behalf.
  • Multi-Agent Systems: Orchestration and communication between multiple AI agents.
  • Agent Development Kit (ADK): Client-side code framework and SDK for building multi-agent solutions.
  • Root Agent: Top-level agent coordinating the overall flow and communicating with the user.
  • Agent Tools: Agents invoked by the root agent to perform specific tasks, with control returning to the root agent.
  • Sub-Agents: Agents that the root agent can transfer the conversation to for more specialized interactions.
  • Callback Context: Mechanisms to execute code before or after agent calls for control and determinism.
  • Tool Context: Short-term memory for storing information related to tool execution.
  • NL2SQL: Natural Language to SQL conversion.
  • BQML: BigQuery Machine Learning.

1. Introduction to AI Agents and Multi-Agent Systems

  • The evolution from simple LLMs and prompting to RAG pipelines, then to adding function calling and tools.
  • AI agents incorporate LLMs, RAG, and tools, but with reasoning, planning, and orchestration capabilities.
  • Multi-agent systems involve orchestrating communication between different agents.
  • AI agents plan, reason, and execute tasks, handling session management and tool execution.

2. Agent Development Kit (ADK) Overview

  • The ADK is a client-side code framework and SDK for building sophisticated multi-agent solutions.
  • It simplifies setting up agents, tools, and multi-agent systems.
  • Out-of-the-box tools include function calls and retrieval tools (e.g., Vertex AI Search, Vertex AI RAG).
  • Custom tools can be created by extending the base tool class.
  • The ADK provides mechanisms to handle artifacts generated during agent workflows.

3. Data Science Agent Architecture

  • Database Agent: Handles database analysis, executes SQL queries against BigQuery, and includes a SQL validator and fixer.
    • Uses two methods for building SQL: Gemini (shown in detail) and JSQL (from Claudia Research).
  • Data Science Agent: Performs NL2PI (Natural Language to Python Interface) workloads, generates plots, and executes Python code using the code interpreter extension.
  • BigQuery ML (BQML) Agent: Supports more sophisticated data science workloads, writes BQML SQL statements for training models and performing inference.
  • Root Agent: Coordinates the overall flow, communicates with the user, and routes queries to the appropriate agents.
    • Provides the Data Science Agent and Database Agent as tools and the BQML Agent as a sub-agent.
  • Data Sources: Connects to BigQuery (can be extended to other databases).

4. Sample Prompt and Agent Interaction

  • Example prompt: "How many rows are in my train table?"
    • The root agent passes the query to the Database Agent.
    • The Database Agent generates and executes the SQL, fetches the data, and returns it to the root agent.
    • The root agent generates an answer for the user.
  • Example prompt: "Can you generate a plot of total sales per country?"
    • The root agent uses the Database Agent to fetch the data.
    • The data is passed to the Data Science Agent to generate the code and execute it.
    • The plot is displayed to the user.

5. Codebase Deep Dive

  • The core agent code resides in the agents/data_science/data_science directory.
  • Key files: agent.py, instructions.py, and tools.py.

5.1. agent.py (Root Agent)

  • Uses the ADK to define and instantiate the root agent.
  • Configuration parameters: model (Gemini 1.5 Pro), agent name, instructions, global instructions, sub-agents, and tools.
  • Sub-agents: BQML Agent (conversation is transferred to this agent).
  • Tools: Database Agent, Data Science Agent, and Load Artifacts.
  • load_artifacts: Python function to manage artifacts generated in the multi-agent flow (e.g., plots).
  • before_agent_call: Callback function to execute code before invoking the agent (e.g., providing BigQuery schema and DDL details).
    • Optimizes efficiency by providing schema information upfront, avoiding unnecessary database queries.

5.2. instructions.py

  • Contains prompts and instructions for the root agent.
  • Instructions on how the agent should consider and use the tools.
  • Defines the workflow (e.g., call the Database Agent first to retrieve data before executing the Data Science Agent).
  • Key reminders and instructions for routing to the BQML sub-agent.

5.3. tools.py

  • Contains the tool definitions for the root agent (Database Agent, Data Science Agent, Load Artifacts).
  • asyn_call to the Database Agent: Pre-processing before invoking the agent tool.
  • asyn_call to the Data Science Agent: Passes the question and data to the Data Science Agent.
  • Tool context: Used to store the output of the tools (e.g., Database Agent output) for later use.

5.4. BigQuery Agent

  • Uses two methods for SQL generation: native Gemini and JSQL.
  • The method can be set in the environment file.
  • Follows the same structure as the root agent (agent.py, instructions.py, tools.py).
  • Includes a before_agent_call callback.

6. Running the Data Science Agent

  • Use the command adk web in the data science directory to spin up the ADK developer front end.
  • The front end allows interacting with the multi-agent system.
  • The UI tracks events, sessions, and artifacts.
  • The evaluation tab allows creating evaluation sets for testing changes.

7. Demo and Examples

  • Example query: "What data do you have?" (answered directly by the root agent using the provided schema).
  • Example query: "What countries exist in my train table?" (invokes the Database Agent).
  • Example query: "Generate a plot of total sales per country" (invokes the Database Agent and then the Data Science Agent).
  • Example query: "I want to train a forecasting model using BQML" (transfers the session to the BQML Agent).

8. Key Takeaways

  • The ADK simplifies building multi-agent systems.
  • The architecture is modular and extensible.
  • The ADK provides built-in tools for artifact management and evaluation.
  • The system can be easily adjusted to specific use cases.

9. Resources

  • GitHub repository for the Agent Development Kit.
  • GitHub repository with agent samples (including the data science agent).
  • Links for the evaluation (Bird SQL eval using JSQL).
  • Video showing how to train a model using BQML to compete in the Kaggle competition.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.