Effective AI Agents Need Data Flywheels, Not The Next Biggest LLM – Sylendran Arunagiri, NVIDIA

AI EngineerAbout 5 min readJun 3, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • AI Agents: Systems that perceive, reason, and act on tasks, learning from user feedback.
  • Data Flywheels: Continuous loops of data processing, curation, model customization, evaluation, and guardrailing for AI agents.
  • Nemo Microservices: An NVIDIA platform with components for each stage of the data flywheel loop (curation, customization, evaluation, guardrails, retrieval).
  • NVIDIA NIM: Optimized inference for models powering agentic use cases.
  • Router Agent: An LLM-powered agent that directs user queries to specific expert agents.
  • Expert Agents: Specialized agents that handle queries within specific domains.
  • Model Drift: Degradation of model performance over time due to changes in data or user behavior.
  • GenAI Ops: Managing the end-to-end generative AI pipeline.

Data Flywheels for AI Agents: Staying Relevant and Helpful

The video focuses on building effective AI agents that remain relevant and helpful over time using data flywheels, rather than solely relying on larger language models (LLMs). It explains what data flywheels are, how they were applied to an internal agent at NVIDIA, lessons learned, and a framework for building data flywheels for various AI agent use cases.

What are AI Agents?

AI agents are defined as systems that can perceive data, reason, and act on tasks. They utilize tools, functions, and external systems to address user queries. A crucial aspect is their ability to learn from user feedback, refining themselves for improved accuracy and usefulness.

The Challenge of Building and Scaling AI Agents

Building and scaling AI agents can be challenging due to:

  • Rapidly changing data: Enterprise data and business intelligence are constantly evolving.
  • Shifting user preferences: Customer needs and expectations change over time.
  • High inference costs: Using larger LLMs increases operational expenses.

Data Flywheels: A Solution

Data flywheels address these challenges by creating a continuous cycle of:

  1. Data Processing and Curation: Refining and organizing enterprise data.
  2. Model Customization: Fine-tuning models for specific tasks.
  3. Evaluation: Assessing model performance and accuracy.
  4. Guardrailing: Ensuring safe and responsible AI interactions.

This cycle leverages production environment data, user feedback, and business intelligence to continuously experiment with and evaluate models. The goal is to identify efficient, smaller models that offer comparable accuracy to larger LLMs but with lower latency, faster inference, and reduced total cost of ownership.

NVIDIA's Nemo Microservices

NVIDIA's Nemo Microservices is an end-to-end platform designed to build powerful agentic and generative AI systems, including data flywheels. It offers various components for each stage of the data flywheel loop:

  • Nemo Curator: Curates high-quality training datasets, including multimodal data.
  • Nemo Customizer: Fine-tunes and customizes models using techniques like LoRa, P-tuning, and full SFT.
  • Nemo Evaluator: Evaluates models using academic and institutional benchmarks, as well as LLM-as-a-judge.
  • Nemo Guardrails: Provides guardrails for privacy, security, and safety.
  • Nemo Retriever: Builds state-of-the-art Retrieval-Augmented Generation (RAG) pipelines.

Nemo Microservices are exposed as simple APIs and can be deployed on-premise, in the cloud, in data centers, or at the edge. NVIDIA provides enterprise-grade stability and support.

Sample Data Flywheel Architecture with Nemo Microservices

The video presents a sample architecture using Nemo Microservices as "Lego pieces" to build a data flywheel. An end-user interacts with an agent's front end, which is guardrailed for safety. The backend uses a model served by NVIDIA NIM for optimized inference.

The data flywheel loop continuously curates data in a Nemo data store and uses Nemo Customizer and Evaluator to retrain and evaluate models. Once a model meets the target accuracy, it can be promoted to power the agentic use case via NVIDIA NIM.

Case Study: NV Info Agent

The video details a real-world case study of applying a data flywheel to NVIDIA's internal employee support agent, NV Info Agent. This agent assists employees with access to enterprise knowledge across various domains (HR, IT, finance, etc.).

Architecture:

  • Users submit queries to the agent, which is guardrailed.
  • A router agent (powered by an LLM) orchestrates multiple expert agents, each specializing in a specific domain.
  • Expert agents use RAG pipelines to fetch relevant information.

A data flywheel is set up to continuously build on user feedback and production inference logs. Subject matter experts provide human-in-the-loop feedback to curate ground truth data. Nemo Customizer and Evaluator are used to evaluate models and promote the most effective model to power the router agent.

Router Agent Deep Dive:

The router agent directs user queries to the appropriate expert agent. The goal is to accurately route queries using faster and more cost-effective LLMs.

Experiment and Results:

  • Initial testing showed a 70B model achieved 96% accuracy in routing, but smaller models (e.g., 8B) had significantly lower accuracy (14%).
  • A feedback form was circulated among NVIDIA employees to collect data on query usefulness.
  • This resulted in 1,224 data points, with 729 satisfactory and 495 unsatisfactory responses.
  • Nemo Evaluator and subject matter experts identified 32 instances of incorrect routing within the unsatisfactory responses.
  • A ground truth dataset of 685 data points was created (60/40 split for training/testing).
  • Fine-tuning smaller models (8B, 1B) with this dataset significantly improved their accuracy.
  • The 8B model matched the 70B model's accuracy after fine-tuning.
  • The 1B model achieved 94% accuracy, only 2% lower than the 70B model.

Benefits:

  • Deploying a 1B model resulted in a 98% reduction in inference costs and a 70x reduction in model size and latency.
  • The data flywheel enables continuous evaluation and fine-tuning as new models are released, allowing for the use of smaller, more efficient models.

Framework for Building Effective Data Flywheels

The video concludes with a framework for building effective data flywheels:

  1. Monitor User Feedback: Collect intuitive user feedback signals (explicit and implicit) to understand model drift and inaccuracies.
  2. Analyze and Attribute Errors: Classify and attribute errors to understand why the agent is behaving in a certain way. Create a ground truth dataset.
  3. Plan: Identify different models, generate synthetic datasets, experiment with fine-tuning, and optimize resource and cost.
  4. Execute: Trigger the data flywheel cycle, track accuracy and latency, monitor performance and production logs, and manage the end-to-end GenAI Ops pipeline.

Conclusion

The key takeaway is that data flywheels are essential for building and maintaining effective AI agents. By continuously learning from data and user feedback, agents can remain relevant, accurate, and cost-effective. NVIDIA's Nemo Microservices and NIM provide the tools and infrastructure to implement data flywheels and optimize AI agent performance. The framework provided offers a structured approach to building and managing data flywheels for various agentic use cases.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.