The Leaderboard Illusion - Gaming the System
By Prompt Engineering
Key Concepts
- Leaderboard Illusion: The systematic issues with benchmarks like LM Arena that lead to a distorted ranking of AI systems.
- LM Arena: A crowdsourced benchmark platform for large language models featuring anonymous, randomized battles.
- ELO Score: A rating system used in LM Arena to rank models based on pairwise comparisons.
- Data Access Disparity: The unequal access to data and feedback mechanisms that benefits proprietary model providers.
- Model Removal: The practice of removing models from a leaderboard, with a disproportionate impact on open-source models.
- Overfitting: Tuning a model specifically for a particular benchmark, leading to poor real-world performance.
- Human Preference Scores: Metrics that measure how well a model aligns with human preferences, which can be manipulated.
- Multiple Comparison Problem: The increased chance of finding a statistically significant result by chance when performing multiple tests.
Systematic Issues with Benchmarks (The Leaderboard Illusion)
The paper "The Leaderboard Illusion" highlights systematic issues with benchmarks, particularly LM Arena, that lead to a distorted playing field. These issues are not limited to LM Arena but represent a broader problem in the LLM community.
LM Arena: A Crowdsourced Benchmark
LM Arena is a leaderboard created in May 2023 that ranks large language models based on anonymous, randomized battles. Users are presented with responses from two LLMs to the same prompt in a blind fashion and asked to choose the better response. This crowdsourced approach has become an industry standard, with Google's CEO even citing LM Arena rankings.
The Llama 4 Incident
The release of Llama 4 Maverick highlighted potential issues with LM Arena. Meta claimed that an experimental chat version of Llama 4 Maverick achieved an ELO score of 1417 on LM Arena. However, the released model performed significantly worse, raising questions about the validity of the leaderboard. The current score of Llama for Maverick is on the 38th rank compared to the second rank when this was released and now instead of that 1400 score it's only scoring about 1,200.
Undisclosed Private Testing Practices
The paper alleges that some providers engage in undisclosed private testing practices, allowing them to test multiple variants before public release and retract scores if desired. The authors claim they were able to test 27 different private models or variants and pick the one that was doing the best and report those ELO scores. This selective disclosure of performance results biases arena scores.
Data Access Disparity and Overfitting
Proprietary model providers benefit disproportionately from free, open community data. They can use this data as a feedback mechanism to fine-tune or fit their models to the LM Arena leaderboard. The paper claims that even small amounts of data can lead to a 112% gain in performance improvement on LM Arena. LM Arena shares 20% of the data back with the model providers.
Model Removal Bias
The paper notes that 66% of silently removed models are open-weight or open-source models, suggesting a bias against open-source models.
Other Benchmark Controversies
Frontier Math Benchmark
A preview version of GPT-3 scored 25% on the Frontier Math benchmark, a significant improvement over the previous best model's 2%. However, OpenAI funded the benchmark's creation and had exclusive access to the hardest problems with solutions, raising questions about potential bias.
ARC AGI Leaderboard
GPT-3 also scored highly on the ARC AGI leaderboard but used some of the training data, which they were not supposed to.
Types of Benchmarks
The video identifies two categories of benchmarks:
- Academic Benchmarks (e.g., GPQA): These benchmarks are often saturated, as model creators may indirectly use data from them.
- Sponsored Benchmarks (e.g., Frontier Math, LM Arena): These benchmarks require careful scrutiny of the sponsors and their potential influence.
The Problem with Human Preference Scores
Alex Albert, head of Claude relations, argues that blindly chasing better human preference scores is a toxic feedback loop, similar to chasing total watch time on social media. It can lead to manipulating users instead of providing genuine value. OpenAI also made GPT40 really bad and had to roll back all of the dates because it was getting really bad.
Community Reactions and LM Arena's Response
Andrej Karpathy's Concerns
Andrej Karpathy expressed suspicion about Gemini models scoring highly on LM Arena despite poor real-world performance. He also noted that Claude models, specialized for coding, rank poorly on LM Arena. He suggests that open router rankings could potentially be a good alternative.
LM Arena's Defense
The LM Arena team acknowledges the concerns but defends its approach. They argue that the leaderboard reflects millions of fresh, real human preferences, which are subjective but important. They are working on statistical methods to decompose human preferences. They also state that they designed their policy to prevent model providers from just reporting the highest score they received during testing and they only published the score for the model they released publicly.
Conclusion
The "Leaderboard Illusion" paper and the ensuing discussion highlight the challenges of creating and maintaining unbiased benchmarks for large language models. Issues such as undisclosed private testing, data access disparity, model removal bias, and the manipulation of human preference scores can distort the rankings and lead to misleading conclusions about model performance. While benchmarks like LM Arena have contributed to the progress of LLMs, it is crucial to address these systematic issues to ensure a fair and accurate evaluation of AI systems.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development