OpenAI Multimodal Analytics Agent: No‑Code Deploy + Repo

Arseny ShatokhinAbout 7 min readOct 30, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Advanced Data Analytics Agent: An AI system designed to analyze data, provide insights, and suggest actions without requiring manual coding for setup or customization.
  • No MCP Servers: A departure from traditional AI agent architectures that relied on Message Queueing Telemetry Transport (MQTT) servers for API integration.
  • IPython Interpreter Tool: A custom tool that allows the AI agent to directly execute Python code, enabling dynamic API requests and data manipulation.
  • Multimodal Tool Outputs: A feature that allows AI agents to return not only text but also images and files from their tool executions, enabling richer interactions and analysis.
  • Vertical Agent: A specialized AI agent designed for a specific industry or use case, requiring minimal customization for deployment.
  • Bundle Pricing: A pricing strategy where clients pay a recurring fee for a package of multiple specialized AI agents, along with support and reporting.
  • Credentials File/API Keys: Information required for the AI agent to authenticate and access external data sources like Google Analytics or Stripe.
  • Tool Output Image/File: Specific data types used by the multimodal tool outputs feature to represent visual or file-based information returned by a tool.

Advanced Data Analytics Agent: Deployment and Architecture

This video introduces a highly advanced, customizable data analytics agent that can be deployed in under 60 seconds without any coding. The agent aims to function like a human data analyst, providing actionable insights from business data.

Demo and Initial Capabilities

The agent's onboarding process involves filling out a single form. When prompted with a request like "how can I improve our website traffic and conversions?", the agent demonstrates the following capabilities:

  • Reasoning: It first analyzes the user's request.
  • Web API Discovery: It searches for relevant APIs to fulfill the request.
  • Data Pulling: It uses a custom IPython interpreter tool to directly pull data from sources like Google Analytics.
  • Dynamic API Request Generation: Crucially, the agent generates API requests on the fly, eliminating the need for pre-configured Message Queueing Telemetry Transport (MQTT) servers.
  • Code Execution: The agent executes code multiple times (over 30 times in the demo) to fetch and process data.
  • Chart Generation and Visualization: It generates charts based on the analyzed data and utilizes OpenAI's new multimodal tool outputs feature to "see" and interpret these charts.
  • Actionable Insights: It provides key findings and actionable recommendations, such as identifying conversion tracking issues and suggesting traffic optimization from YouTube, all backed by calculated statistics.

The entire setup and analysis for this demo took approximately 60 seconds, with the cost for the analysis being around three cents.

Architectural Breakthroughs

The development of this agent was driven by two key realizations:

  1. Ditching MCP Servers: The presenter argues that MCP servers are a "horrible idea" and not a great abstraction for AI agents.

    • Problems with MCP Servers:
      • Data is dumped directly into the agent's context window, which is limited.
      • Large Language Models (LLMs) are primarily trained on text, not numbers, making direct numerical data analysis challenging.
    • The New Approach:
      • A simple, ~50-line Python tool using an IPython shell is created.
      • The agent is prompted to interact with APIs directly by generating Python code for requests.
    • Benefits of Direct API Interaction:
      • Flexibility: The agent is not limited to pre-defined MCP servers and can generate requests to any API.
      • Leverages LLM Training: Agents can already work with API data as it's similar to their training data.
      • Complex Data Handling: The agent can fetch data from multiple sources simultaneously, perform complex transformations, and then provide insights, mimicking a human analyst.
      • Data Management: The agent can save data to the file system and only read key statistics, preventing context window overload and improving accuracy.
    • Limitations of Direct API Interaction:
      • Increased Latency: Generating API requests each time takes longer.
      • Higher Token Consumption: More tokens are used due to the dynamic generation process.
      • Reduced Reliability: The logic for API requests is not hard-coded, leading to a chance of incorrect API usage.
    • Justification for Trade-offs: For data fetching (not data updates), the trade-offs are considered valid. Latency and token consumption are deemed manageable, and the risk of incorrect requests is low when only reading data.
  2. Multimodal Tool Outputs Feature: This feature significantly enhances agent capabilities by allowing tools to return images and files, not just text.

    • Impact: Unlocks new use cases in areas like software development (e.g., agents can "see" websites to fix layouts).
    • Future Agents: The team plans to release a QA tester agent and an ad creator agent that leverage this feature.

Agent Architecture and Workflow

The agent is designed with simplicity, featuring only four tools. The core functionality lies within the read_image tool, which enables the agent to interpret generated charts.

The final workflow is as follows:

  1. Read API Docs: The agent reads API documentation directly.
  2. Create API Requests: It generates requests to fetch data.
  3. Analyze Data: The fetched data is analyzed.
  4. Create Charts (if needed): Visualizations are generated.
  5. Read Charts: The agent interprets the generated charts using the read_image tool.
  6. Provide Insights: Actionable insights are delivered to the user.

Pricing and Deployment Strategy

The recommended pricing strategy for AI agents is based on the value they create, aiming for a 10x return on investment for the client, which is a common practice in B2B SaaS.

  • Vertical Agent Pricing: While the cost of individual agents is low, they are designed as "vertical agents," meaning they require minimal customization for deployment to multiple clients. Clients pay a recurring fee for ongoing use.
  • Bundle Pricing: The current strategy is to offer bundle pricing at around $2,500 per month, which includes 3-5 specialized, customized agents, onboarding, Slack support, and monthly reports. This approach is encouraged for others to replicate.
  • Customization: The presenter plans to release a future video on creating custom vertical agents.

Deployment Process

The agent is deployed through a platform called "agency."

  1. Access Marketplace: Navigate to the platform's marketplace.
  2. Select Agent: Choose the data analytics agent.
  3. Deploy and Fill Form: Click "deploy" and complete an onboarding form with key questions:
    • Business Overview: General information about the business.
    • Business Goals: Objectives the agent should help achieve.
    • Analytics APIs: List of platforms used (e.g., Google Analytics, Stripe API). For each, specify its purpose and whether a credentials file or API key is used.
    • Credentials File/API Key: Upload a credentials file or provide API keys. The presenter suggests using ChatGPT to get detailed guides for clients on obtaining these.
    • Reasoning Effort: (Implied, not explicitly detailed in the demo)
    • Output Format: Desired format for the agent's output.
  4. Save and Build: Save the form, and the agent will automatically build and deploy within approximately one minute.
  5. Integration: The deployed agent can be integrated into custom GPTs, Slack, or Zapier. Future integrations with Chrome Jobs and Notion are planned.
  6. Web Application Deployment: Agents can be deployed into web applications by selecting the agent, customizing the app, and hitting deploy.
  7. Platform Benefits: The platform is free to start, and deployed agents automatically receive future updates. New fields in the onboarding form will trigger notifications for users to fill them out.

Future Plans and Improvements

  • Scrape API Docs: Scrape documentation for common analytics APIs to enhance the agent's reliability, making it comparable to MCP servers.
  • Expert Training: Hire a professional data analyst to train the agent, as the current workflow could be improved.
  • Memory for Mistakes: Implement a memory feature for the agent to recall and avoid repeating incorrect API requests, especially with less common APIs.

Technical Details: Multimodal Tool Outputs

For advanced users, the multimodal tool outputs feature is explained:

  • Implementation: Tools need to return a specific type: tool_output_image or tool_output_file_content.
  • Helper Methods:
    • tool_output_image_from_path: Converts a local image to the required type.
    • tool_output_file_from_url: Converts a file from a URL.
  • Combining Outputs: Multiple output types can be combined, allowing for multiple images with annotations.
  • Open Source: The code for the analytics agent and the multimodal feature is available in an open-source repository, allowing users to copy, fork, and create their own versions.
  • Workshops: The presenter will host workshops on building agents from scratch and productizing them.

Conclusion

The presented data analytics agent represents a significant advancement in AI agent capabilities, offering a powerful, customizable, and cost-effective solution for businesses. The shift away from MCP servers towards direct API interaction, coupled with the multimodal tool outputs feature, enables more flexible, accurate, and versatile data analysis. The platform's ease of deployment and pricing model further democratize access to advanced AI tools.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.