2025 is the Year of Evals! Just like 2024, and 2023, and … — John Dickerson, CEO Mozilla AI

AI EngineerAbout 6 min readAug 7, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • AI/ML Monitoring and Evaluation
  • Agentic Systems
  • Generative AI (GenAI)
  • Large Language Models (LLMs)
  • Return on Investment (ROI)
  • Downstream Business KPIs
  • LLM as a Judge Paradigm
  • Multi-Agent Systems
  • Open Source AI

Main Topics and Key Points

Thesis: 2025 - The Year of Evals

  • The speaker posits that 2025 will be the year that AI/ML evaluation becomes a top priority for businesses.
  • AI/ML monitoring and evaluation are presented as two sides of the same coin. Monitoring requires measurement, which is the core of evaluation.
  • Three concurrent factors are driving this shift:
    • CEOs, CFOs, and CISOs now understand AI due to the accessibility of tools like ChatGPT.
    • Budget freezes in enterprises created an opportunity for AI projects to receive funding.
    • AI systems are now acting autonomously, necessitating rigorous evaluation.

The Rise of Agentic Systems

  • Agentic systems are defined as systems that perceive, learn, abstract, generalize, reason, and act.
  • These systems are increasingly being deployed in enterprises and SMBs, leading to greater complexity and risk.
  • The speaker emphasizes that the need for evaluation becomes critical when systems start making decisions and taking actions on behalf of humans.

Historical Context: Pre-ChatGPT Era

  • Before November 30, 2022 (the launch of ChatGPT), ML monitoring existed, but it was often disconnected from downstream business KPIs.
  • Selling AI/ML monitoring solutions was challenging because the need for evaluation was not obvious to the entire C-suite.
  • While many companies were founded in the AI/ML monitoring space (H2O, Algorithmia, Seldon, Y Labs, Aporeia, Arise, Arthur, Galileo, Fiddler, Protect AI, Snowflake, Databricks, DataDog, SageMaker, Vertex, etc.), it was rarely a top priority for decision-makers.
  • The speaker mentions Jamie Dimon's annual report from JPMC in April 2022, noting that the AI spend, while significant, was still relatively small compared to the overall budget.

The Impact of Economic Conditions and ChatGPT

  • Economic uncertainty in late 2022 led to budget freezes in many enterprises.
  • ChatGPT's launch created a "pet project" opportunity for AI, as CEOs and CFOs were impressed by its capabilities.
  • Discretionary budgets were unlocked specifically for GenAI projects.

The Evolution from Science Projects to Production

  • 2023 was characterized by GenAI science projects within enterprises.
  • In 2024, GenAI-based applications started going into production, primarily internally.
  • This led to questions from the C-suite regarding ROI, governance, risk, compliance, and brand optics, driving the need for evaluation.
  • 2025 is seeing scale-up, with increased revenue for frontier model providers and AI evaluation companies.
  • IT budgets for 2024 and beyond are earmarked specifically for AI applications.

The Role of the C-Suite

  • The speaker highlights the importance of aligning the C-suite on the need for AI evaluation.
  • CEO: Understands the capabilities of generative models and agentic systems and is comfortable allocating budget.
  • CFO: Focuses on the impact to the bottom line and requires quantitative evaluation to justify investments.
  • CISO: Sees AI as a security risk and opportunity, leading to investments in guardrail products.
  • CIO: Remains on board and supports the need for evaluation.
  • CTO: Seeks standards and data-driven decision-making, which evaluation provides.

Multi-Agent Systems Monitoring

  • Evaluation companies are shifting towards multi-agent systems monitoring, recognizing the need to monitor the entire system, not just individual models.
  • The speaker references an article in The Information that leaked revenue numbers for AI evaluation startups, noting that those numbers are now outdated due to recent growth.

Mozilla AI's Contribution

  • Mozilla AI offers an open-source "light LLM" for multi-agent systems called "Any Agent."

Important Examples, Case Studies, or Real-World Applications Discussed

  • JPMC: Jamie Dimon's annual report is cited as an example of early AI investment, though the speaker notes the relative smallness of the investment.
  • Internal Chat Applications and Hiring Tools: Mentioned as examples of GenAI-based applications going into production in 2024.
  • Discounted Cash Flow (DCF) Analysis: Used as an example of a complex task performed by multi-agent systems that requires domain expertise and human validation.

Step-by-Step Processes, Methodologies, or Frameworks Explained

  • Agent Definition: The speaker provides a definition of an agent, emphasizing the ability to perceive, learn, abstract, generalize, reason, and act.

Key Arguments or Perspectives Presented, with Their Supporting Evidence

  • Evaluation is Essential for AI Adoption: The speaker argues that evaluation is crucial for quantifying risk, demonstrating ROI, and ensuring responsible AI deployment.
  • The C-Suite's Role in Driving Evaluation: The speaker emphasizes that the C-suite's understanding and support are essential for making evaluation a priority.
  • Multi-Agent Systems Require Comprehensive Monitoring: The speaker argues that monitoring individual models is insufficient and that the entire multi-agent system must be monitored.

Notable Quotes or Significant Statements with Proper Attribution

  • "I see AI/ML monitoring and evaluation as two sides of the same sword or ruler."
  • "2025 we all hear it. It's year of the agent. I no longer is a question mark needed here. It's clearly the the year of the agent."

Technical Terms, Concepts, or Specialized Vocabulary with Brief Explanations

  • AI/ML Monitoring: The process of tracking the performance and behavior of AI/ML models in production.
  • AI/ML Evaluation: The process of assessing the quality, accuracy, and fairness of AI/ML models.
  • Agentic Systems: AI systems that can perceive, learn, reason, and act autonomously.
  • Generative AI (GenAI): AI models that can generate new content, such as text, images, or code.
  • Large Language Models (LLMs): A type of AI model that is trained on vast amounts of text data and can generate human-like text.
  • Return on Investment (ROI): A measure of the profitability of an investment.
  • Downstream Business KPIs: Key performance indicators that are directly impacted by the performance of AI/ML models.
  • LLM as a Judge Paradigm: Using LLMs to evaluate the performance of other AI models.
  • Multi-Agent Systems: Systems composed of multiple interacting AI agents.

Logical Connections Between Different Sections and Ideas

  • The speaker begins by establishing the thesis that 2025 will be the year of evals.
  • The speaker then provides historical context, explaining why evaluation has not been a top priority in the past.
  • The speaker then discusses the impact of economic conditions and ChatGPT, which created an opportunity for AI projects to receive funding.
  • The speaker then explains how AI projects are moving from science projects to production, driving the need for evaluation.
  • The speaker then highlights the role of the C-suite in driving evaluation.
  • The speaker then discusses the shift towards multi-agent systems monitoring.
  • The speaker concludes by mentioning Mozilla AI's contribution to the open-source AI community.

Any Data, Research Findings, or Statistics Mentioned

  • Jamie Dimon's annual report from JPMC, which showed that the company had invested $100 million in AI/ML from 2017 to 2021.
  • Leaked revenue numbers for AI evaluation startups from The Information.
  • A leaked spreadsheet from Merkor showing hourly rates for experts hired by large companies to validate AI systems.
  • A paper in iClear discussing biases that LLMs as judges have versus humans.

Brief Synthesis/Conclusion of the Main Takeaways

The speaker argues that 2025 will be the year that AI/ML evaluation becomes a top priority for businesses due to the confluence of increased understanding of AI by the C-suite, the availability of funding for AI projects, and the rise of agentic systems. The speaker emphasizes the importance of aligning the C-suite on the need for evaluation and highlights the shift towards multi-agent systems monitoring. The speaker concludes by mentioning Mozilla AI's contribution to the open-source AI community.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.