How We Built Zeta2: Training an Edit Prediction Model in Production — Ben Kunkle, Zed

AI EngineerAbout 4 min readMay 31, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • Edit Prediction: A specialized AI task where a model predicts the next code edit based on cursor position, surrounding code, type definitions, and diagnostics.
  • Distillation: The process of using a large "frontier" model (teacher) to generate training data for a smaller, faster "student" model (Zed 2).
  • Settled Data: A technique of capturing code snapshots after a user has finished editing a specific region, used to refine training data.
  • Levenshtein Distance: A string metric used to measure the difference between two sequences (used here to compare model predictions against "settled" code).
  • Reversal Ratio: A heuristic metric tracking how often a model suggests an edit that immediately undoes the user's previous input.
  • JSONL (JSON Lines): The data format used for the training pipeline, where each line is a self-contained JSON object, allowing for fluid field manipulation across pipeline stages.

1. The Edit Prediction Pipeline

The goal of Zed 2 is to provide low-latency code suggestions that run on every keystroke. Because the task is highly specialized, the team utilizes a small, fine-tuned model rather than a general-purpose LLM.

  • Data Collection: The system captures opt-in production data, including snapshots of code, cursor context, variable definitions, and diagnostic errors.
  • Distillation Process:
    1. Teacher Generation: Frontier models are prompted to predict edits based on the collected context.
    2. Repair Step: To handle the inconsistency of frontier models (which may provide different answers for the same input), the team uses a secondary "repair" prompt. If a prediction fails static heuristics (e.g., ignoring boundaries or undoing user work), it is sent back to a frontier model to be corrected.
    3. Prompt Formatting: This stage is experiment-specific, determining which features (like diagnostics or edit history) are included in the final training prompt.

2. Leveraging "Settled Data"

A significant challenge in training is determining what constitutes a "good" edit. The team uses "settled data"—snapshots taken after a user stops editing a region for 10 seconds.

  • Filtering Noise: Because user intent can change, settled data is inherently noisy. To filter this, the team generates 10 predictions from the student model and uses Levenshtein distance to see if any are close to the settled state.
  • The "Goldilocks" Zone:
    • Too far: Considered noise/irrelevant.
    • Too close: Obvious, trivial completions.
    • The Middle: The ideal training data, often containing new functions or patterns outside the student model’s original training cutoff.

3. Evaluation Methodologies

The team employs both offline and online evaluations to ensure model quality:

  • Offline Evals: Conducted on a held-out test set. They use Delta Car F (an n-gram comparison tool) and track the reversal ratio. They generate three distinct teacher predictions per input to account for the fact that there is rarely only one "correct" way to write code.
  • Online Evals: Once deployed, models are tested in production using a traffic-splitting dashboard (e.g., 15% of traffic). They monitor acceptance rates, latency, and diagnostic error counts (comparing error rates before and after the prediction).

4. Key Arguments and Perspectives

  • Efficiency over Generality: Ben Kunkel argues that for an editor, a small, specialized model is superior to a large one because it must execute on every keystroke with minimal latency.
  • Iterative Pipeline: The team emphasizes a modular pipeline where data is stored in JSONL format. This allows them to cache intermediate steps and run multiple experiments by simply adding or reformatting fields without re-running the entire distillation process.
  • Self-Correction: By using the student model to evaluate its own potential training data (instead of relying solely on expensive frontier models), the team can scale their training data volume significantly while keeping costs low.

5. Synthesis and Conclusion

The Zed 2 training methodology represents a sophisticated approach to "learning from the user." By combining distillation from frontier models with self-supervised filtering of "settled" production code, the team has created a high-performance, specialized model. The pipeline is designed for rapid experimentation, utilizing a modular JSONL-based architecture that allows for quick iteration on features like diagnostic inclusion and edit history, ultimately resulting in a model that is rapidly approaching the quality of the frontier models used to teach it.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.