Key Concepts
- Data Filtering: Selecting a subset of raw data that is similar to a target dataset.
- Deduplication: Removing duplicate or near-duplicate entries from a dataset.
- N-gram Model: A statistical language model that predicts the probability of a word given the preceding n-1 words.
- Kneser-Ney Smoothing: A technique used in n-gram models to handle unseen n-grams by interpolating or backing off to lower-order n-grams.
- Fasttext: A library for text classification and representation learning, known for its speed and efficiency.
- Importance Sampling: A technique used to estimate properties of a target distribution by sampling from a different proposal distribution and weighting the samples accordingly.
- Hash Functions: Functions that map items to hash values, used for efficient data lookup and deduplication.
- Bloom Filter: A space-efficient probabilistic data structure used to test whether an element is a member of a set.
- Jaccard Similarity: A measure of similarity between two sets, defined as the ratio of the size of their intersection to the size of their union.
- MinHash: A hash function with the property that the probability of a hash collision is equal to the Jaccard similarity between the input sets.
- Locality Sensitive Hashing (LSH): A technique used to find near neighbors in high-dimensional spaces by hashing similar items into the same buckets with high probability.
Filtering Algorithms
N-gram Model with Kneser-Ney Smoothing
- Concept: Train an n-gram model on a target dataset and use it to score raw data, selecting the top-scoring documents.
- Details:
- Uses Kneser-Ney smoothing to handle sparse counts and unseen n-grams.
- Implementation often uses KenLM, an open-source toolkit.
- Fitting the model involves counting n-grams and normalizing.
- Maximum Likelihood Estimation (MLE) is the starting point, estimating conditional probabilities like P(word_n | word_n-1, ..., word_1).
- Example: CCNet used this approach, sorting paragraphs by increasing perplexity and keeping the top 1/3 for the Llama dataset.
- Pros: Fast and simple.
- Cons: Crude and only considers local contexts.
Fasttext Classifier
- Concept: Train a linear classifier to distinguish between target data and raw data.
- Details:
- Maps the vocabulary space into a smaller hidden dimension to reduce the number of parameters.
- Uses asynchronous Stochastic Gradient Descent (SGD) for optimization.
- Extends to n-grams by hashing n-grams into a fixed number of bins to handle the unbounded number of bigrams.
- Often used with K=2 classes (good/bad).
- Trade-off: Using larger models like BERT or Llama for classification may be slower and less efficient than training the model directly.
Data Selection with Importance Sampling
- Concept: Build an importance weight estimator to sample documents from a raw dataset that are similar to a target dataset.
- Details:
- Uses importance re-sampling to compensate for sampling from a proposal distribution (raw data) instead of the target distribution.
- Fits distributions p to target data Dp and q to raw data Dq.
- Uses hash n-grams to estimate the probability of each hashed n-gram.
- Calculates importance weights as the ratio of the target distribution to the raw distribution.
- Advantage: More principled approach for diversity as it tries to match the distribution.
- Improvement: Can be improved by using better models instead of linear models based on n-grams.
Applications of Filtering
Language Identification
- Goal: Find text of a specific language.
- Method: Use a pre-trained fasttext language identification model.
- Example: DOMA uses fasttext and keeps pages with English probability > 0.5.
- Considerations: Short sentences are less reliable, low-resource languages are tough, and dialects can be misclassified.
Quality Filtering
- Goal: Filter out low-quality data.
- Methods:
- GPT-3 trained a quality classifier using high-quality sources as positive samples and Common Crawl as negative samples.
- Llama used pages referenced by Wikipedia as positives and sampled Common Crawl as negatives.
- phi-1 used GPT-4 to determine the educational value of Python code and trained a random forest classifier on the resulting data.
- Key Idea: Use a strong language model to synthesize or filter down a lower quality dataset to get a target dataset.
Toxicity Filtering
- Goal: Filter out toxic or NSFW content.
- Method: Train fasttext classifiers on datasets like Jigsaw Toxic Comments.
- Example: DOMA trains two fasttext classifiers, one for hate and one for NSFW.
Deduplication Algorithms
Exact Deduplication
- Concept: Remove exact duplicates from a dataset.
- Method:
- Compute a mapping from hash value to the list of items with that hash.
- Keep one item from each group.
- Implementation: Can be implemented using MapReduce or hash tables.
- Example: C4 uses exact match deduplication at the three-sentence span level.
- Pros: Simple, clear, and high precision.
- Cons: Doesn't work for near-duplicates.
Bloom Filters
- Concept: A space-efficient probabilistic data structure for approximate set membership testing.
- Details:
- Uses a bit array and multiple hash functions.
- To add an item, hash it with k hash functions and set the corresponding bits in the array.
- To query an item, hash it with k hash functions and check if all corresponding bits are set.
- Can have false positives but no false negatives.
- False positive rate can be controlled by adjusting the number of bins and hash functions.
- Formula: The optimal value of K is scaling as order M over N.
- Example: Dolma sets the false positive rate to 10^-15 and uses a bloom filter to do their exact deduplication at the paragraph level.
Near-Deduplication with MinHash and LSH
- Concept: Remove near-duplicates from a dataset based on Jaccard similarity.
- Details:
- Uses MinHash to estimate the Jaccard similarity between two sets.
- Uses Locality Sensitive Hashing (LSH) to find near neighbors in linear time.
- LSH involves breaking up hash functions into b bands of r hash functions.
- Two items collide if for some band, all its hash functions return the same value.
- Tuning: Increasing r sharpens the curve and moves things to the right, and increasing b will shift the curve to the left.
- Example: One paper uses b=20 and r=450, resulting in a threshold of 0.99.
Conclusion
The lecture provides a comprehensive overview of data filtering and deduplication techniques used in training language models. It covers various algorithms, including n-gram models, fasttext classifiers, importance sampling, bloom filters, and MinHash LSH. The lecture emphasizes the importance of algorithmic tools and spending time with the data to build intuitions about what works and what doesn't. The key takeaway is that data filtering and deduplication are essential steps in preparing high-quality training data for language models, leading to improved performance and efficiency.
AI summaries can miss context or contain errors. Check important details against the original video.