Key Concepts
- Brand Search Optimization (BSO): Automating search engine optimization workflows for a brand.
- Agent Development Kit (ADK): Google's framework for building AI agents.
- Sub-agents: Specialized agents within a larger multi-agent system (Keyword Finding Agent, Search Results Agent, Comparison Root Agent).
- Web Browser Agent: An agent that uses a web browser (Selenium, WebDriver) to interact with websites, enabling dynamic content loading and user interaction.
- BigQuery: Google's fully managed, serverless data warehouse.
- Evaluation (Eval): Process of assessing the performance and robustness of an agent using predefined criteria and sample conversations.
- Selenium: A portable framework for testing web applications.
- Gemini Model: Google's large language model used for various tasks within the agents.
Brand Search Optimization Agent Overview
The video presents a brand search optimization agent built using Google's Agent Development Kit (ADK). This agent automates the SEO workflow for a brand by optimizing titles based on a comparison of top search results for relevant keywords. The agent uses a multi-agent architecture, incorporating sub-agents for keyword finding, search result retrieval, and comparison.
Agent Architecture and Flow
- Keyword Finding Agent:
- Purpose: Identifies relevant keywords for a given brand based on its titles, descriptions, and attributes stored in a BigQuery table.
- Tools: BigQuery client to retrieve data.
- Process:
- Retrieves brand data (title, description, attributes) from BigQuery using a SQL query.
- Uses the Gemini model to extract keywords from the retrieved data.
- Ranks and groups keywords based on criteria like genericness and similarity.
- Output: A ranked list of keywords.
- Search Results Agent:
- Purpose: Retrieves the titles of top search results for the identified keywords from a specified website (e.g., Google Shopping).
- Tools: Web browser agent (Selenium, WebDriver) with tools for:
go_to_url: Navigates to a specified URL.take_screenshot: Captures a screenshot of the current page.find_element_with_text: Locates an element on the page based on its text content.click_with_text: Clicks on an element based on its text content.enter_text_into_element: Enters text into a specified element (e.g., a search bar).scroll_down: Scrolls down the page.get_page_source: Retrieves the HTML source code of the page.analyze_web_page_and_determine_action: Determines the series of steps needed to achieve an end goal.
- Process:
- Takes a keyword as input.
- Navigates to the specified website.
- Enters the keyword into the search bar.
- Retrieves the HTML source code of the search results page.
- Extracts the titles of the top search results.
- Output: A list of titles from the top search results.
- Comparison Root Agent:
- Purpose: Compares the brand's titles with the titles of the top search results and generates a report with recommendations for optimization.
- Sub-agents:
- Generator Agent: Creates an initial comparison report.
- Critique Agent: Improves the comparison report based on predefined policies and rules.
- Process:
- The Generator Agent compares the brand's titles with the top search result titles.
- The Critique Agent refines the comparison, incorporating company-specific policies and business rules.
- The Generator and Critique agents loop until a satisfactory comparison report is generated.
- Output: A report with recommendations for optimizing the brand's titles.
Why Use a Computer Vision Agent?
The video highlights the advantages of using a computer vision agent (web browser agent) over traditional web scraping APIs:
- Dynamic Content Loading: Modern websites often load data dynamically using JavaScript. Web browser agents can render these dynamic elements, while web scraping APIs may only retrieve the initial HTML source code.
- User Interaction: Some websites require user interaction (e.g., clicking buttons, scrolling down) to access data. Web browser agents can mimic these interactions.
- Anti-Scraping Measures: Web browser agents can better handle anti-scraping measures implemented by websites.
- Data Extraction from Non-HTML Elements: Web browser agents can extract data from elements that are not directly present in the HTML source code.
- Complex Website Navigation: Web browser agents can navigate complex website structures by mimicking real user behavior.
Code Implementation Details
- Root Agent (agent.py): The main agent that routes requests to the sub-agents. It uses prompting to define the workflow steps. An alternative approach is to use the
SequentialAgentclass fromgoogle.kate.agents. - Keyword Finding Agent: Uses a tool to query BigQuery and retrieve brand data. The tool needs to be replaced if BigQuery is not used.
- Search Results Agent: Uses Selenium and WebDriver to interact with websites. It includes tools for navigating, interacting with elements, and retrieving the page source.
- Comparison Agent: Uses a generator and critique agent to create and refine the comparison report. The looping between these agents can be implemented using prompt instructions or the
LoopAgentclass from ADK.
Evaluation and Testing
- Evaluation (Eval): The ADK provides an evaluation tool to assess the performance of the agent. Evaluation data is generated by saving agent sessions as JSON files. The
eval.shscript runs the evaluation using the saved data and a configuration file. - Unit Tests: Unit tests are included for the BigQuery data tool, using
magic_mockto mock the BigQuery data. - Example Interaction: An example interaction is provided in a markdown format to demonstrate the end-to-end flow of the agent.
Running the Agent
- Environment Setup:
- Set up a BigQuery table with sample data (instructions in the
readme). - Copy the
.env.examplefile to.envand configure the environment variables. - Set
DISABLE_WEB_DRIVERto0for the web app and1for evaluation and unit testing.
- Set up a BigQuery table with sample data (instructions in the
- Start the Web App:
- Navigate to the
brand_search_optimizationfolder. - Run the command to start the ADK web app.
- Navigate to the
- Interact with the Agent:
- Open the web browser and navigate to the ADK web app.
- Select the
brand_search_optimizationagent. - Provide the brand name and follow the prompts to guide the agent through the workflow.
Evaluation Process
- Create Evaluation Set:
- Open the web browser and go to the evaluation tab.
- Create an evaluation set and give it a name.
- Save the session as an eval set (JSON file).
- Run Evaluation Script:
- Use the
eval.shscript to run the evaluation. - The script uses the ADK eval command-line tool with the folder name, data set file, and config path.
- Use the
Conclusion
The brand search optimization agent demonstrates the power of ADK for building AI agents that automate complex tasks. By combining sub-agents with specialized tools and leveraging the capabilities of large language models, the agent can effectively optimize brand titles for search engines. The evaluation and testing tools provided by ADK ensure the robustness and reliability of the agent. The video provides a comprehensive overview of the agent's architecture, implementation, and usage, making it a valuable resource for developers interested in building AI-powered SEO solutions.
AI summaries can miss context or contain errors. Check important details against the original video.





