Key Concepts
- Observability and Evals: Rigor and excellence in AI product development stem from observability and good evaluations.
- Notion AI: A suite of AI-powered features integrated into the Notion workspace, focusing on content generation, automation, and search.
- Brain Trust: A platform for AI evaluation, experimentation, and observability, enabling teams to iterate on AI models and prompts effectively.
- LLM as a Judge: Using large language models to evaluate the quality and correctness of AI outputs, either with a single prompt for all data or specific prompts for each data element.
- Offline Evals: Structured tests run on predefined data sets to evaluate and iterate on prompts and models before deployment.
- Online Evals: Real-time tracing and scoring of AI outputs in production, allowing for monitoring performance, diagnosing problems, and capturing user feedback.
- Human-in-the-Loop: Incorporating human review and user feedback into the AI development process to improve quality, reliability, and alignment with user needs.
- Remote Evals: Extending the Brain Trust playground by connecting to a local development environment, enabling the evaluation of complex tasks with intermediate code steps and external system integrations.
Notion AI and its Evolution
- Early Access and Content Generation: Notion AI launched before Chat GPT, focusing on content generation with the AI Writer.
- Autofill and Database Integration: Introduced autofill for database properties, enabling AI to act across databases and perform tasks like translation.
- Natural RAG Solution: Developed a natural retrieval-augmented generation (RAG) solution with a full data platform and collaboration features, offering Q&A to free users.
- Universal Search and Attachments: Launched universal search across apps and the ability to search over attachments.
- AI Meeting Notes and Deep Research: Recent launches include AI meeting notes with transcription and summaries, and a deep research tool with agentic capabilities.
Challenges in AI Evaluation
- Large Datasets: Notion AI deals with exceptionally large datasets, requiring scalable evaluation solutions.
- Overwhelmed Human Evaluators: Human evaluators were initially overwhelmed by the volume of data and the complexity of parsing prompts.
- Quality vs. Quantity: Emphasized that quality is more important than quantity in human labeling, particularly for fine-tuning and iteration.
Iteration Cycle with Brain Trust
- Decide on an Improvement: Identify an area for improvement, such as launching a Jira connector in universal search.
- Curate Targeted Datasets: Create handcrafted datasets from logs and prototypes, focusing on structured data.
- Tie to Scoring Functions: Develop scoring functions specific to the product, including LLM as a judge and heuristic-based functions.
- Inspect Results: Analyze the evaluation results and identify areas for improvement.
- Iterate: Continuously refine the data, scoring functions, and prompts based on the evaluation results.
LLM as a Judge System
- Two Components:
- Single Prompt Judge: A single prompt judges everything in the dataset (e.g., "Is this information concise?").
- Specific Prompt per Element: Each element in the dataset has a particular prompt tailored to its expected output (e.g., "This answer should be in Japanese, bullets should be formatted this way").
- Benefits of Specific Prompts:
- Captures expected value more accurately than Levenshtein distance.
- Allows for more up-to-date golden sets by specifying rules rather than fixed outputs.
- Facilitates quick assessment of new models by running evals and identifying prompts that perform well.
Outcomes and Impact
- Critical to Iteration Flow: Brain Trust is critical to Notion AI's iteration flow and is considered part of their intellectual property.
- Skyrocketed Product Quality: Observability provided by Brain Trust has significantly improved AI product quality.
- Multilingual Support: Rigorous evaluation metrics enable building AI products that work for a majority of non-English speakers.
Brain Trust Core Concepts
- Prompt Engineering: Rapidly iterate on prompts using a playground environment.
- Automated Evals: Use the SDK to kick off automated evaluations and generate scores.
- Observability: Monitor production performance, capture user feedback, and diagnose problems.
Components of an Eval
- Task: The code or prompt being evaluated, ranging from simple LLM calls to complex agentic workflows.
- Dataset: A set of real-world examples or test cases used to evaluate the task.
- Score: The logic behind the evaluation, outputting a score from 0 to 100, using LLM as a judge or heuristic functions.
Mental Models for Evals
- Offline Evals: Iterating on prompts and models using structured tests and predefined datasets.
- Online Evals: Monitoring live production applications, getting scores from real outputs, and capturing user feedback.
Score Types
- LLM as a Judge: Subjective, non-deterministic, qualitative assessment of AI outputs.
- Heuristic Score: Exact, deterministic, objective evaluation based on code.
Brain Trust UI: Playgrounds vs. Experiments
- Playgrounds: For quick iteration and ephemeral testing.
- Experiments: For more traditional experiments, triggered via the SDK or CI pipeline, providing a historical view of performance over time.
Activity 1: Hands-on with the Brain Trust Platform
- Objective: To create evals, run experiments, and understand the Brain Trust platform.
- Steps:
- Install Node, Git, and TypeScript.
- Create a Brain Trust organization and project.
- Configure the AI provider (OpenAI).
- Clone the provided Git repository.
- Copy the
env.local.examplefile toenv.localand replace the API keys. - Run
pnpm installto install dependencies and create resources. - Explore the created prompts, scores, and datasets in the Brain Trust platform.
- Create an initial playground and load the prompts and scores.
- Run the eval and analyze the scores.
- Create an experiment to track scores over time.
Running Evals via the SDK
- Process:
- Define assets (prompts, scores) as code.
- Run the
brain trust evalcommand via the CLI or CI pipeline. - View the experiments in the Brain Trust platform.
Logging and Online Scoring
- Why Logging:
- Measure the quality of live traffic.
- Debug and troubleshoot issues.
- Close the feedback loop.
- How to Log:
- Wrap the LLM client using
wrapOpenAI. - Trace arbitrary functions using
tracedecorator orwrapTrace. - Use
span.logfor granular control and metadata.
- Wrap the LLM client using
- Online Scoring:
- Use scores from offline evals in production.
- Set up online scoring rules for specific spans.
- Start with a low sampling rate and increase as trust in metrics grows.
- Views:
- Create custom views to filter and sort logs based on specific criteria.
Human-in-the-Loop
- Why Human-in-the-Loop:
- Catch hallucinations.
- Establish ground truth.
- Capture user feedback.
- Types of Human-in-the-Loop:
- Human Review: Annotators manually label data sets or logs.
- User Feedback: Incorporate user feedback into the development process.
- Process:
- Enable user feedback in the application.
- Create human review scores in the Brain Trust configuration.
- Assign human evaluators to review logs and provide feedback.
- Incorporate feedback into data sets and prompts.
Remote Evals
- Problem: The playground cannot handle complex tasks with intermediate code steps or external system integrations.
- Solution: Remote evals allow you to expose a local development environment to the playground.
- How it Works:
- Brain Trust sends an eval request to your remote server.
- Your server runs the task and score logic locally.
- The server returns the outputs, scores, and metadata to Brain Trust.
- Benefits:
- Evaluate complex tasks with custom tooling and intermediate code steps.
- Bridge the gap between technical teams and non-technical users.
Conclusion
The presentation provides a comprehensive overview of AI evaluation and observability using Brain Trust, emphasizing the importance of rigorous testing, continuous iteration, and incorporating human feedback. It covers various aspects, from setting up evaluations in the UI and SDK to logging production data and implementing human-in-the-loop workflows. The introduction of remote evals extends the capabilities of the Brain Trust platform, enabling the evaluation of complex tasks with custom tooling and external system integrations.
AI summaries can miss context or contain errors. Check important details against the original video.