Vector Search Benchmark[eting] - Philipp Krenn, Elastic

AI EngineerAbout 5 min readJun 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Benchmarking, vector search, performance evaluation, use case selection, data bias, read-only vs. read-write workloads, filtering, precision/recall, statistical manipulation, reproducibility, nightly benchmarks, slow boiling frog problem, custom benchmarks, flaw analysis.

Benchmarking Challenges and Pitfalls

The Problem with Existing Benchmarks

  • Many benchmarks claim "X is faster than Y," but the specific products represented by X and Y can vary widely, leading to contradictory results.
  • Benchmarks are often difficult to interpret and compare due to varying methodologies and configurations.
  • "Glossy charts" and visually appealing presentations can sometimes mask underlying flaws or biases in the data.
  • Quote: "I think at this point we have benchmarks for every single vendor being both faster and slower than all of their competitors."

Use Case Selection and Data Bias

  • Companies often design benchmarks that favor their own systems, creating scenarios where they perform well while competitors struggle.
  • Benchmarking datasets can be cherry-picked to highlight specific strengths, leading to generalized claims of superiority that don't hold true across all use cases.
  • Example: A company might run various scenarios and only publicize the one where their system outperforms competitors, ignoring other results.
  • Comic Analogy: The speaker references a comic where two systems are compared under "similar conditions," but one is clearly favored, illustrating the common practice of biased benchmarking.

Read-Only vs. Read-Write Workloads

  • Most vector search benchmarks focus on read-only datasets because they are easier to reproduce and compare.
  • Real-world workloads are often read-write, making read-only benchmarks less representative.
  • Factors like read-write ratio and data update frequency are often ignored in benchmarks, leading to inaccurate performance assessments.

Filtering in Vector Search

  • Filtering can paradoxically slow down vector search, especially with HNSW (Hierarchical Navigable Small World) indexes.
  • Unlike traditional databases where filtering reduces the dataset size and speeds up ranking, vector search may require examining more candidates before filtering, increasing latency.
  • Companies can exploit this behavior by tuning scenarios to favor systems with specific filtering optimizations.

Versioning and Updates

  • Benchmarks often compare the latest version of a company's own software against older versions of competitors' software.
  • This practice can artificially inflate performance gains, as competitors may have released significant improvements since the benchmarked version.
  • Keeping up with the changes of all competitors is difficult, but using outdated versions introduces bias.

Implicit Biases and Default Settings

  • Developers' familiarity with their own systems can lead to unintentional biases in benchmark design.
  • Default settings, such as shard size, memory allocation, and instance configuration, can be optimized for one system but detrimental to others.
  • These biases can arise even without malicious intent, simply from a lack of comprehensive benchmarking across different configurations.

Cheating and Statistical Manipulation

  • "Cheating" can involve manipulating parameters or implementations to achieve better benchmark results, even at the expense of accuracy or quality.
  • Example: The Volkswagen emissions scandal is cited as an example of creative ways to look good in benchmarks.
  • In vector search, precision and recall are often overlooked, leading to comparisons of systems with vastly different result quality.
  • Statistical manipulation can involve highlighting a single, highly optimized use case to create the illusion of overall superiority.
  • Example: A system might be significantly faster in one specific query type (e.g., ascending vs. descending sorting) and then use this outlier to inflate overall performance statistics.

Reproducibility Issues

  • Benchmarks that lack detailed documentation or reproducible code make it difficult for others to verify the results.
  • The absence of transparency undermines the credibility of the benchmark and makes it harder to identify potential flaws.

Improving Benchmarks

Automation and Reproducibility

  • Benchmarks should be automated and run regularly to track performance changes over time.
  • Example: The speaker's company uses "nightly benchmarks" to monitor performance and identify regressions.
  • Automated benchmarks help prevent the "slow boiling frog problem," where gradual performance degradation goes unnoticed.
  • Analogy: The "slow boiling frog problem" refers to the idea that small, incremental changes can lead to significant negative consequences if not monitored closely.

Custom Benchmarks

  • Users should create their own benchmarks tailored to their specific use cases, data characteristics, and performance requirements.
  • Existing benchmarks are unlikely to perfectly match a user's unique needs, making custom benchmarks essential for accurate evaluation.
  • Factors to consider include data size, data structure, read-write ratio, query patterns, acceptable latency, and hardware configuration.
  • Tool Mentioned: "Rally" is a tool used internally to create and run custom benchmarks against their own system.

Flaw Analysis and Learning

  • Instead of dismissing flawed benchmarks entirely, users should analyze them to identify potential strengths and weaknesses of different systems.
  • Even a biased benchmark can reveal insights into a system's sweet spot or the scenarios it is designed to excel in.
  • Understanding these strengths and weaknesses can inform better decision-making and optimization strategies.

Conclusion

The speaker emphasizes the importance of critical thinking when evaluating benchmarks, highlighting the numerous ways in which they can be misleading or biased. He advocates for automated, reproducible, and custom benchmarks tailored to specific use cases. Even flawed benchmarks can provide valuable insights if analyzed carefully. The key takeaway is that users should not blindly trust benchmarks but instead conduct their own evaluations to make informed decisions.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.