Stanford CS336 Language Modeling from Scratch | Spring 2026 | Lecture 14: Data

Stanford OnlineAbout 4 min readMay 29, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • Data Pipeline: The end-to-end process of transforming, filtering, deduplicating, and mixing raw data for model training.
  • Rule-based vs. Model-based Processing: Using heuristics (rules) for speed versus using machine learning models for intelligent data extraction.
  • Filtering: The process of selecting high-quality subsets from raw data using classifiers (e.g., FastText).
  • Deduplication: Removing exact or near-duplicate content to improve training efficiency and prevent memorization.
  • MinHash & LSH (Locality Sensitive Hashing): Algorithmic techniques to identify near-duplicates in linear time.
  • Data Mixing: Determining the optimal distribution of different data sources to maximize model performance.
  • Synthetic Data: Data generated by "teacher" models to train "student" models, particularly for post-training/reasoning tasks.

1. Data Transformation

Raw data from sources like Common Crawl is typically HTML, PDF, or code repositories.

  • HTML Processing: Primarily rule-based to remove boilerplate (ads, navigation, footers). It is inherently lossy because it linearizes hierarchical/visual structures into token sequences.
  • PDF Processing: PDFs are high-value but difficult to process. They often require OCR (Optical Character Recognition) or Vision Language Models (VLMs). They are often truncated in web crawls, necessitating re-crawling.
  • Key Insight: Rule-based processing is preferred for its speed, though model-based interventions are becoming more viable for complex content extraction.

2. Data Filtering

Filtering is essential for "compute-poor" environments to avoid wasting resources on low-quality content.

  • Framework: Define a "target" (high-quality) set and a "raw" (large) set. Train a classifier (e.g., FastText) to distinguish between them.
  • Applications:
    • Language Identification: Using off-the-shelf FastText models to filter by language.
    • Domain-Specific: Projects like Open Math Text use a combination of rules (LaTeX detection) and classifiers to curate high-quality math corpora, outperforming models trained on 20x more unfiltered data.
    • Toxicity: Using annotated datasets (e.g., Jigsaw) to filter out harmful content.
  • Scaling: The optimal filtering threshold depends on the total training token count. If training for longer, one can tolerate lower-quality data; if training for shorter, high-quality data is critical.

3. Deduplication

Deduplication prevents "wasting GPUs" on redundant content and mitigates memorization/privacy risks.

  • Exact Deduplication: Removing identical strings or document spans.
  • Near-Deduplication: Identifying documents with high Jaccard similarity (e.g., >0.99).
  • MinHash & LSH:
    • MinHash: A hashing technique where the probability of a collision equals the Jaccard similarity of two sets.
    • LSH: Uses bands and rows of hash functions to create a "phase transition" in collision probability, allowing for efficient identification of near-duplicates at scale.

4. Data Mixing

Mixing involves assigning weights to different data sources (e.g., books, code, web text).

  • The Epoch Problem: Naive mixing can lead to over-training on small, high-quality datasets (e.g., training 50 epochs on a small high-quality set while barely touching a large low-quality set).
  • Methodologies:
    • Unimax: Samples sources uniformly but enforces a hard cap on the number of epochs per source.
    • Regression-based Mixing (RegMix): Training small proxy models on various mixtures to fit a regression function, then optimizing that function to find the best mixture for large-scale training.
    • Simulated Epoching: Downsampling small-scale experiments to mimic the data scarcity/epoching behavior of large-scale runs.

5. Post-Training Data

Post-training focuses on task-specific capabilities (reasoning, coding, agentic behavior).

  • Synthetic Data Generation: Using strong "teacher" models to generate responses for prompts.
  • Coding Agents: Projects like SweetZero demonstrate that models can learn agentic coding tasks (e.g., GitHub PRs) without needing live execution feedback, by distilling trajectories from larger models.
  • Key Finding: Better models are not always better teachers; sometimes smaller, specialized models (e.g., QWQ 32B) outperform frontier models in specific reasoning tasks.

Synthesis and Conclusion

Data work is described as "grungy" and highly domain-specific. The core takeaway is that data quality and composition are as important as model architecture. For most practitioners, the path to a high-performing model involves:

  1. Aggressive filtering using lightweight classifiers.
  2. Rigorous deduplication using LSH to save compute.
  3. Principled mixing that accounts for epoch counts to prevent overfitting.
  4. Strategic use of synthetic data for post-training to imbue models with specific reasoning or agentic skills.

Notable Quote: "If you have infinite compute, you don't need to filter... but realistically, everyone has to filter."

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.