YouTube Recommendation Systems with Large Language Models: A Deep Dive
Key Concepts:
- Large Language Models (LLMs) for Recommendations
- Semantic ID (SID)
- Generative Retrieval
- Continued Pre-training
- Tokenization of Videos
- User Context
- Personalized Recommendations
- Offline Recommendation Tables
- LLM-Augmented Recommendations
1. Introduction: The Underhyped Potential of LLMs in Recommendations
- The speaker argues that the application of LLMs to recommendation systems is a larger consumer application than search, despite receiving less attention.
- YouTube recommendations drive a large majority of watch time across various surfaces like Home, Watch Next, and Shorts.
- The core problem is learning a function that takes user and context as input and outputs relevant recommendations.
- YouTube leverages user demographics, watch history, engagement metrics (comments, subscriptions), and other contextual data to generate video recommendations.
2. LRM: Adapting Gemini for YouTube Recommendations
- YouTube has developed a system called LRM (Large Recommender Model) by adapting the Gemini LLM specifically for video recommendations.
- The process involves:
- Starting with a base Gemini checkpoint.
- Continued pre-training on YouTube-specific data to create a unified YouTube-specific checkpoint.
- Aligning the model for various recommendation tasks like retrieval and ranking.
- Creating smaller, custom versions of the model for different recommendation surfaces.
- LRM is currently in production for retrieval and under experimentation for ranking.
3. Semantic ID: Tokenizing Videos for LLMs
- Problem: LLMs require tokenization of input. Representing videos directly is inefficient due to context length limitations.
- Solution: Semantic ID (SID), a method to create semantically meaningful tokens for videos.
- Process:
- Extract features from a video: title, description, transcript, audio, and video frame data.
- Embed these features into a multi-dimensional embedding space.
- Quantize the embedding using Residual Quantization with Vector Expansion (RQVE) to assign each video a token.
- Significance: SID creates a "language of YouTube videos" where tokens represent semantic concepts (e.g., music, gaming, sports). Videos sharing similar topics share prefix tokens, while unique identifiers differentiate them.
- SID moves away from hash-based tokenization to a semantically meaningful representation.
4. Continued Pre-training: Linking Text and Video Tokens
- The goal is to teach the LLM to understand both English and the new "YouTube language" (SID tokens).
- Two-step process:
- Linking Text and SID: Training the model to associate video titles, creators, and topics with their corresponding SID tokens. For example, prompting the model with "This video has title XYZ" and training it to output the correct title.
- Understanding Watch Sequences: Training the model to understand relationships between videos based on user engagement data (watch history). For example, prompting the model with "A user has watched videos A, B, C, D" and masking some videos, training the model to predict the masked videos.
- This pre-training results in a model that can reason across English and YouTube videos.
- Example: The model can infer that a video about AI is likely interesting to technology fans based on the SID representation and the user's past watch history of tennis, F1, and math videos.
5. Generative Retrieval: Personalized Recommendations with LRM
- Construct a prompt for each user including:
- Demographic information (age, gender, location, device).
- Context video (the video the user is currently watching).
- Watch history and engagement data.
- Have the model decode video recommendations as SID tokens.
- Benefits:
- Generates unique recommendations, especially for users with limited data.
- Can identify connections that traditional systems miss.
- Example: When a user watches an Olympics highlight video, LRM can recommend related women's races based on user demographics and watch history, whereas the production system would only recommend other men's track races.
6. Challenges and Solutions: Serving LRM at Scale
- Challenge: LRM is computationally expensive to serve, especially at YouTube's scale.
- Solutions:
- Cost Reduction: Significant (95%+) reduction in TPU serving costs to enable production deployment.
- Offline Recommendation Tables: Generating unpersonalized recommendations offline using LRM. This involves removing personalized aspects from the prompt and building a table of candidate videos for each video. This allows for simple lookup during serving, bypassing the need for real-time LRM inference.
- Even unpersonalized recommendations from LRM are differentiated due to the model's pre-training on a large checkpoint.
7. Key Differences Between Training LLMs and LLM-Based Recommendation Systems
- Vocabulary and Corpus Size: YouTube's video library is vastly larger and more dynamic than the vocabulary of a typical LLM.
- Freshness: Recommending new videos quickly is crucial on YouTube, whereas LLMs can tolerate some staleness.
- Scale: Serving LLMs to billions of daily active users requires smaller, more efficient models.
- Continuous Pre-training: LRM requires continuous pre-training on the order of days or hours, unlike classical LLMs which are pre-trained less frequently.
8. LLM and Recommender Systems Recipe
- Tokenize Your Content: Create domain-specific tokens representing the essence of your content. This can be achieved by extracting features, building embeddings, and quantizing them.
- Adapt the LLM: Link English and your domain language. Find training tasks that help the model reason across English and the new tokens. This creates a bilingual LLM that can speak both natural language and your domain-specific language.
- Prompt with User Information: Construct personalized prompts with user demographics, activity, and actions. Train task-specific or surface-specific models. This results in a generative recommendation system on top of an LLM.
9. Future Directions: Interactive and Generative Recommendations
- Currently, LLMs augment recommendations invisibly, enhancing quality without direct user interaction.
- The future involves users interacting with the recommendation system in natural language, steering recommendations towards their goals.
- Recommenders will be able to explain their recommendations, and users can align the system to their preferences expressed in natural language.
- The lines between search and recommendation will blur.
- Recommendation and generative content will converge, leading to personalized versions of content and even the creation of entirely new content tailored to individual users.
10. Conclusion: The Transformative Potential of LLMs in Recommendations
LLMs are poised to revolutionize recommendation systems, offering significant improvements in personalization, relevance, and user experience. While challenges remain in terms of scalability and computational cost, ongoing research and development are paving the way for a future where recommendations are more interactive, intelligent, and tailored to individual needs. The application of LLMs to recommendations is a significant and underhyped area with the potential to transform how users discover and engage with content.
AI summaries can miss context or contain errors. Check important details against the original video.