Key Concepts
- Evaluation: Assessing the quality and capabilities of language models.
- Benchmarks: Standardized tests used to evaluate language models (e.g., MMLU, GPQA, HellaSwag).
- Perplexity: A metric measuring how well a language model predicts a dataset; lower perplexity indicates better performance.
- Few-shot/Zero-shot Learning: Prompting strategies where models are given a few examples or no examples, respectively, before being asked to perform a task.
- Chain of Thought (CoT): A prompting technique where the model is encouraged to explain its reasoning process step-by-step.
- Train-Test Overlap (Contamination): The presence of data from the test set in the training set, leading to inflated performance metrics.
- Instruction Following: The ability of a language model to perform tasks based on natural language instructions.
- Chatbot Arena: A platform where users interact with two anonymous models and vote on which response is better, used to rank language models.
- Agents: Systems that combine a language model with tools and a planning mechanism to perform complex tasks.
- Safety Benchmarks: Evaluations designed to assess the potential for harmful outputs from language models.
- Jailbreaking: Techniques used to bypass the safety mechanisms of a language model.
Evaluation: An Overview
The lecture focuses on the complexities of evaluating language models, highlighting the "evaluation crisis" due to the proliferation of benchmarks and the uncertainty about their true meaning. While mechanically evaluation seems simple (prompting a model, getting responses, computing metrics), it profoundly influences language model development.
The Purpose of Evaluation
The speaker emphasizes that there is no single "true" evaluation. The purpose of evaluation depends on the question being asked:
- Users/Companies: Making purchase decisions (e.g., choosing between Claude, Grok, Gemini, or GPT-4).
- Researchers: Understanding the raw capabilities of models and scientific progress in AI.
- Policymakers/Businesses: Objectively understanding the benefits and harms of models.
- Model Developers: Getting feedback to improve models.
A Framework for Evaluation
A simple framework for thinking about evaluation is presented:
- Inputs (Prompts):
- Where do the prompts come from?
- Which use cases are covered?
- Do they represent edge cases?
- Are they adapted to the model?
- Calling the Language Model:
- Few-shot, zero-shot, chain of thought prompting.
- Tool use (arithmetic, retrieval-augmented generation).
- Evaluating the language model vs. the entire system (agent).
- Outputs:
- Are reference outputs clean and error-free?
- What metrics are used (e.g., pass@1, pass@10 for code generation)?
- How is cost factored in?
- How are different types of errors weighted?
- How is open-ended generation evaluated?
- Interpreting Results:
- What does a score of 91% mean?
- Has the model truly learned generalization?
- Is the object of evaluation the model, the system, or the method?
Perplexity: A Deep Dive
Perplexity measures whether a language model assigns high probability to a dataset. It's calculated against a validation set.
- Historical Context: In the 2010s, perplexity was the primary metric for language modeling research, with datasets like Penn Treebank and WikiText-103 being commonly used. Google's work showed that scaling up model size and designing the architecture correctly could dramatically reduce perplexity.
- GPT-2 and Beyond: GPT-2 shifted the focus to downstream task accuracy, but perplexity remains useful. GPT-2 trained on 40GB of text from websites linked from Reddit and evaluated directly on standard perplexity benchmarks without fine-tuning.
- Usefulness: Perplexity is smoother than downstream task accuracy and is used in scaling laws. It's also "universal" in the sense that it considers every token in the dataset.
- Caveats: Perplexity evaluations require trusting the language model provider to generate valid probabilities.
- Perplexity Maximalism: The idea that minimizing perplexity forces the model to be as close as possible to the true distribution, potentially leading to AGI.
- Related Tasks: Cloze tasks (e.g., LAMBADA) and common sense reasoning tasks (e.g., HellaSwag) are related to perplexity but have been largely saturated by language models.
Standard Knowledge Benchmarks
MMLU (Massive Multitask Language Understanding)
- A canonical benchmark with 57 subjects, consisting of multiple-choice questions collected from the web.
- Originally designed to evaluate base models, but now used for instruction-tuned models.
- The speaker notes that MMLU is more about testing knowledge than language understanding.
- Few-shot prompting is used, but the choice of examples matters.
- Concerns exist about overfitting to MMLU.
MMLU Pro
- An improved version of MMLU that removes noisy questions and increases the number of answer choices to 10.
- Chain of thought prompting is commonly used.
GPQA (Google-Proof Question Answering)
- Focuses on hard, PhD-level questions.
- Questions are validated by experts and non-experts.
- Designed to be difficult to answer even with Google search.
- GPT-4 achieves 39% accuracy, while GPT-3 achieves 75%.
Humanity's Last Exam (HLE)
- A multimodal benchmark with multiple-choice and short-answer questions.
- A prize pool was created to encourage the creation of problems.
- Frontier language models are used to reject questions that are too easy.
- GPT-3 achieves 20% accuracy.
Instruction Following Benchmarks
Chatbot Arena
- A popular benchmark where users interact with two anonymous models and vote on which response is better.
- ELO scores are computed to rank the models.
- Dynamic and incorporates fresh data.
- Concerns exist about privileged access and less-than-ideal evaluation protocols.
IFEval
- Tests the ability of a language model to follow constraints (e.g., answer with a specific number of sentences, use certain words).
- Constraints can be automatically verified.
- Evaluates constraint following, not necessarily the semantics of the response.
AlpacaEval
- Computes a win rate against a particular model as judged by a language model (GPT-4).
- Biased because it asks GPT-4 to evaluate against its own generation.
- Correlated with Chatbot Arena.
WildBench
- Utterances come from human-bot conversations.
- Uses LM as a judge with a checklist.
- Also correlated with Chatbot Arena.
Agent Benchmarks
Agent benchmarks evaluate systems that combine a language model with tools and a planning mechanism.
SweetBench
- Given a codebase and a GitHub issue description, the goal is to submit a PR that makes the unit tests pass.
SideBench
- For cyber security, the agent has access to a server and must hack into it to retrieve a secret key.
- The agent architecture involves planning, generating commands, and iterating.
- Accuracies are still fairly low (up to 20%).
MLE Bench
- Involves 75 Kaggle competitions where the agent must write code, train a model, debug, and submit.
- Accuracies are also low (sub-20%).
ARC (Abstraction and Reasoning Corpus) AGI Challenge
- Focuses on reasoning and creativity rather than knowledge.
- Tasks involve pattern recognition and filling in missing elements.
- Language models have traditionally performed poorly on these tasks.
- GPT-3 is now doing pretty well on this task.
Safety Benchmarks
Safety benchmarks assess the potential for harmful outputs from language models.
Harm Bench
- Identifies 510 harmful behaviors and prompts a language model to see if it will follow the instructions.
AirBench
- Anchors safety in regulatory frameworks and company policies.
Jailbreaking
- Techniques used to bypass the safety mechanisms of a language model.
- A paper developed a procedure to optimize prompts to bypass safety, even transferring to GPT-4.
Pre-Deployment Testing
- Safety institutes are working with model developers to run safety evaluations before release.
- Safety is strongly contextual and depends on law, politics, and social norms.
- Capabilities and propensity (to cause harm) are important considerations.
Realism and Validity
- Standardized benchmarks are often far from real-world use cases.
- Real-world traffic can be spammy.
- Asking prompts (where the user doesn't know the answer) are more realistic than quizzing prompts.
- Anthropic uses language models to analyze real-world data.
- Med-ELM focuses on realistic use cases in medicine.
- Realism and privacy are often at odds.
- Train-test overlap is a major concern.
- Data set quality is also important.
Conclusion
The lecture concludes by emphasizing the importance of defining the rules of the game and thinking about the purpose of evaluation. The speaker highlights the shift from evaluating methods to evaluating systems and encourages algorithmic innovation.
AI summaries can miss context or contain errors. Check important details against the original video.





