Fuzzing in the GenAI Era — Leonard Tang, Haize Labs

AI EngineerAbout 7 min readAug 24, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • LLM Slop: Subjective and unstructured output from Large Language Models.
  • Hazing (Fuzz Testing): Large-scale optimization and simulation to pressure test AI systems before deployment.
  • Last Mile Problem in AI: Difficulty in making demo-ready AI products robust, enterprise-grade, and reliable for production.
  • Brittleness (Lipschitz Discontinuity): Sensitivity of GenAI systems to slight variations in input, leading to wildly different outputs.
  • Standard Evals: Traditional AI evaluation using static, finite golden datasets, which are insufficient for GenAI.
  • Coverage: The extent to which an evaluation dataset represents the full range of possible inputs.
  • Judging: The process of translating subjective criteria into quantitative metrics for evaluating AI outputs.
  • LLM as a Judge: Using an LLM to evaluate the output of another AI system based on a given prompt or rubric.
  • Scaling Judge Time Compute: Increasing the computational resources and complexity of the judging process to improve accuracy and reliability.
  • Verdict: A library for building agent-based judging systems using principles from scalable oversight.
  • Scalable Oversight: A subfield of AI safety focused on using weaker models to audit, correct, and steer stronger models.
  • GRPO (Generative Reward Policy Optimization): A reinforcement learning technique for training reward models to provide more fine-grained and tailored feedback.
  • Fuzzing: Generating variations of inputs to test the robustness of an AI system.
  • Adversarial Testing: Emulating malicious actors to identify vulnerabilities and weaknesses in AI systems.
  • Discrete Optimization: A mathematical approach to finding the best solution from a finite set of possibilities.
  • Rubric Fanout: A judging architecture that proposes individual unit tests and criteria for each data point, critiques them, self-verifies the critiques, and aggregates the results.

Why Haze? The Problem of AI Unreliability

The speaker introduces Haze as a solution to the problem of unreliable AI systems. AI systems are hard to trust and need to be pressure-tested before deployment. Haze uses large-scale optimization, simulation, and search to determine if a system will behave as expected before it goes into production. The "last mile problem" in AI is highlighted: it's easy to create demo-ready AI, but difficult to make it robust and reliable for enterprise use. Despite promises of autonomy and agency, the lack of trust and reliability is holding back full GenAI adoption.

The Inadequacy of Standard Evals

The speaker argues that traditional evaluation methods, which rely on static golden datasets and human-generated ground truth, are insufficient for GenAI due to the "brittleness" or "Lipschitz discontinuity" of these systems. This brittleness means that small changes in input can lead to drastically different outputs. Examples of this brittleness include AI customer support hallucinating information, AI chatbots giving harmful advice, and AI systems making absurd errors.

Standard evals are insufficient in two ways:

  1. Coverage: Static datasets only reflect performance on that specific data, not the broader input space.
  2. Quality Measurement: It's difficult to create a good measure of similarity between AI outputs and ground truth. The ideal would be a human subject matter expert who can translate their sensitivity into a quantitative metric, but this is a core challenge in AI reward modeling. Current methods like exact match, classifiers, and LLM as a judge have their own quirks and undesirable properties.

Haze: Fuzz Testing in the AI Era

Haze is presented as a form of fuzz testing for AI. It involves:

  1. Simulating large-scale stimuli to send to the AI application.
  2. Getting responses from the AI application.
  3. Judging, analyzing, and scoring the outputs.
  4. Using the scores to guide the next round of search.

This process is repeated iteratively until bugs and corner cases are discovered. If no issues are found within the search budget, the system is considered ready for production. Executing this process is technically difficult, especially in scoring the output and generating the input stimuli.

Judging: Translating Subjective Criteria into Quantitative Metrics

The speaker discusses the challenges of "judging" AI outputs, which involves translating subjective criteria into quantitative metrics. Using LLM as a judge is a common approach, but it has several failure modes:

  • Hallucinations: LLMs are prone to hallucinating information.
  • Instability: Even with good articulation of criteria, LLMs may not operationalize them well.
  • Uncalibration: The scale of an LLM's judgment (e.g., 1 to 5) may not align with human understanding.
  • Bias: LLM judgments can be influenced by input order, context, and rubric variations.

The key question is how to QA the judge itself to ensure it's a gold standard metric.

Scaling Judge Time Compute: Two Approaches

The speaker presents two approaches to improving the judging process by "scaling judge time compute":

  1. Training Reasoning Models from Scratch: Train models specifically for the evaluation task without inductive biases.
  2. Building Agent-Based Judges: Use off-the-shelf LLMs with strong inductive priors to build agents that perform the judging task.

The speaker introduces "Verdict," a library for building agent-based judging systems. Verdict incorporates principles from scalable oversight, such as having weaker LLMs debate each other or self-verify their responses. This approach can create powerful, cheap, and low-latency judging systems. A plot is shown demonstrating that Verdict, powered by a GP40 mini backbone, can outperform larger models like 01, GP4, and 3.5 sonnets on expert QA verification at a fraction of the cost and latency.

RL for Judging: GRPO Tuning

The speaker discusses using reinforcement learning (RL), specifically GRPO tuning, to improve LLM judges. RL can address two issues with standard LLM judges:

  1. Lack of coherent rationales explaining the judgment.
  2. Lack of fine-grained, tailored criteria for specific tasks.

The speaker mentions a paper from Deepseek called SPCT (Self-Principled Critique Tuning), which involves having an LLM propose data-point-specific criteria and critique the data point against those criteria. Experiments using a variant of this technique to GRPO train smaller models (600M and 1.7B parameters) achieved competitive performance on the reward bench task, rivaling larger models like Cloud3 Opus and GP4 mini.

Input Generation: Fuzzing and Adversarial Testing

The speaker discusses two approaches to generating inputs for testing AI systems:

  1. Fuzzing: Generating variations of customer happy paths to test the system under reasonable, in-distribution inputs.
  2. Adversarial Testing: Emulating malicious actors to identify vulnerabilities and weaknesses, such as prompt injections and jailbreaks.

Fuzzing in the AI sense is more structured and optimization-driven than in classical security. Due to the vastness of the natural language input space, brute-force search is impossible. The task is treated as a discrete optimization problem, where the goal is to find inputs that break the AI application according to the judge's score. Various optimization algorithms can be used, including gradient-based methods, tree search, and latent space search.

Case Studies

The speaker presents several case studies:

  1. Largest Bank in Hungary: Haze was used to test a loan calculation AI application, revealing prompt injections, jailbreaks, and unexpected corner cases that were not accounted for in the code of conduct.
  2. Fortune 500 Bank: Haze is being used to test outbound debt collection with voice agents, introducing variance in the audio signal (background noise, static, frequency changes).
  3. Voice Agent Company: Using Verdict to scale up subjective human annotation resulted in a 38% increase in ground truth human agreement compared to using internal ops teams. This was achieved using a "rubric fanout" architecture, which proposes individual unit tests and criteria for each data point, critiques them, self-verifies the critiques, and aggregates the results.

Conclusion

Hazing is crucial for building reliable AI systems, especially in regulated industries. The speaker emphasizes the company's rapid growth and need for more team members.

Q&A

The speaker confirms that the hazing input can be both multi-shot and single-shot, and supports persistent conversations for voice applications.

Main Takeaways

  • The "last mile problem" in AI is the difficulty in making AI systems robust and reliable for production.
  • Traditional evaluation methods are insufficient for GenAI due to the brittleness of these systems.
  • Haze, or fuzz testing, is a crucial approach to pressure-testing AI systems before deployment.
  • Judging AI outputs is challenging, and LLM as a judge has limitations.
  • Scaling judge time compute, through agent-based judges or RL tuning, can improve the accuracy and reliability of the judging process.
  • Fuzzing and adversarial testing are important for generating inputs to test AI systems.
  • Hazing is particularly important for regulated industries like finance and healthcare.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.