Key Concepts
- Foundation Model (FM) for Recommendation
- User Representation Learning
- Transformer Architecture
- Semi-Supervised Learning
- Tokenization of User Interaction History
- ID Embedding vs. Semantic Content Information
- Cold-Start Problem
- Multi-Task Learning
- Scaling Laws in Recommendation Systems
- Multi-Token Prediction
- Multi-Layer Representation
- Generative Retrieval
- Prompt Tuning
Diverse Recommendation Needs at Netflix
- Netflix faces diverse recommendation needs across various dimensions:
- Rows: Different genres (comedies, action), new releases, trending titles, Netflix-exclusive content.
- Items: Movies, TV shows, games, live streaming, and expanding content types.
- Pages: Homepage, search page, kids' homepage, mobile feed (linear layout).
- Traditionally, this diversity led to many specialized models developed independently.
- These models had overlaps, leading to duplication in label and feature engineering.
- Example: Feature engineering for user interaction history involved similar factual data but with slight variations, making maintenance difficult.
Challenges with Traditional Approach
- The traditional approach of developing specialized models was not scalable.
- Spinning up new models for each use case was not manageable.
- Limited leverage and reuse of components impacted innovation velocity.
Foundation Model Approach: Hypothesis
- The key question was: Can we centralize the learning of user representation in one place?
- The answer was yes, based on the hypothesis of using a foundation model (FM) based on transformer architecture.
- Two key hypotheses:
- Scaling Law: Semi-supervised learning can improve personalization through scaling, similar to LLMs.
- High Leverage: Integrating the FM into all systems can simultaneously improve all downstream canvas-facing models.
Data and Training
Data Cleaning and Tokenization
- Tokenization decisions have a profound impact on model quality, similar to LLMs.
- Instead of language tokens (one ID), each token represents a user interaction event with multiple facets/fields.
- Careful consideration is needed for the granularity of tokenization and its trade-off with the context window.
- Different tokenization versions can be used for pre-training and fine-tuning against specific applications.
Model Architecture
- The model architecture consists of:
- Event Representation: Encoding "when," "where," and "what" of an event.
- When: Time encoding.
- Where: Physical location, device, canvas, page.
- What: Target entity/title, interaction type, duration.
- Embedding/Feature Transformation: Combining ID embedding learning with semantic content information to address the cold-start problem.
- Transformer Layer: Using the hidden state output as the user representation.
- Considerations: Stability of user representation, aggregation across time/sequence/layers, explicit adaptation based on downstream objective.
- Objective/Loss Function: Using multiple sequences to represent the output.
- Targets: Entity IDs, action type, entity metadata (genre, language), action duration, device, time of next play.
- Casting the problem as multi-task learning, multi-head/hierarchical prediction, or using targets as weights/rewards/masks.
- Event Representation: Encoding "when," "where," and "what" of an event.
Scaling Laws
- Scaling laws apply to recommendation systems.
- Observed gains from scaling up from millions to billions of model parameters.
- Stopped scaling due to latency cost requirements, but further scaling is possible with distillation.
Learnings from LLMs
- Borrowed learnings from LLMs:
- Top Multi-Token Prediction: Forces the model to be less myopic, more robust to serving time shift, and targets long-term user satisfaction.
- Multi-Layer Representation: Techniques like layer-wise supervision, self-distillation, and multi-layer output aggregation improve user representation.
- Long Context Window Handling: Techniques from truncated sliding window to sparse attention and progressively training longer sequences improve efficiency and maximize learning.
Serving and Applications
Algo Stack Transformation
- Before FM: Many data sources, features, and models developed independently for each canvas/application.
- With FM: Consolidated data and representation layers (user and content), with application models built on top of FM.
FM Utilization Patterns
- Three main approaches:
- Subgraph Integration: Integrating FM as a subgraph within the downstream model, replacing existing sequence transformer towers.
- Embedding Push: Pushing out content and member embeddings to a centralized embedding store for wider use cases (personalization, analytics).
- Concerns: Refresh frequency and stability of member embeddings.
- Model Extraction and Fine-Tuning: Extracting and fine-tuning the FM for specific applications, including distillation for latency requirements.
Results
- Incorporating FM into various applications resulted in:
- High leverage in AB test wins.
- Infrastructure consolidation.
- Scalable solution with improved quality.
- Faster innovation velocity.
Current Directions
- Future directions:
- Universal Representation for Heterogeneous Entities: Semantic ID embedding to cover diverse content types.
- Generative Retrieval for Collection Recommendation: Generating recommendations at inference time, handling business rules and diversity in the decoding process.
- Faster Adaptation through Prompt Tuning: Training soft tokens for inference-time swapping to prompt the FM to behave differently.
Q&A Highlights
- Beyond Recommendation: Expanding to capture user taste from both on and off the platform.
- Graph Models and Reinforcement Learning: Using graph models for knowledge graph and content space coverage, and reinforcement learning for sparse reward guidance in collection generation.
- Granularity of Embeddings: Currently using metadata over the video, but trending towards more granular clip-level or view-level embeddings.
Synthesis/Conclusion
Netflix's adoption of a foundation model for recommendations represents a significant shift towards centralized learning and improved scalability. By leveraging techniques from LLMs and adapting them to the unique challenges of recommendation systems, Netflix has achieved significant gains in model quality, infrastructure consolidation, and innovation velocity. The ongoing efforts to develop universal representations, generative retrieval methods, and prompt tuning techniques promise further advancements in personalization and content discovery.
AI summaries can miss context or contain errors. Check important details against the original video.