Key Concepts
- Semantic IDs: Content-aware item identifiers that encode item characteristics, addressing cold-start and sparsity issues.
- Data Augmentation: Using Language Models (LLMs) to generate synthetic data and labels, enriching datasets for search and recommendation systems.
- Unified Models: Consolidating multiple recommendation, search, and advertising models into a single, versatile model to reduce redundancy and improve efficiency.
- Co-start: The problem of recommending new or unpopular items with limited interaction data.
- Co-star Velocity: The rate at which new items reach a certain threshold of views or interactions.
- Trainable Multimodal Semantic IDs: A method of combining static content embeddings (visual, text, audio) with dynamic user behavior to create semantic IDs.
- Exploratory Search: Helping users discover new categories or items beyond their usual searches.
- Query Recommendation System: Recommending related search queries to users to facilitate exploratory search.
- Multi-task Learning: Training a single model to perform multiple related tasks simultaneously.
- Unified Contextual Ranker: A unified model that ranks items for search, similar item recommendations, and pre-query recommendations.
- Quality Vector: A vector representing the quality of an item based on ratings, freshness, and conversion rate.
- RQVE (Residual Quantized Vector Embedding): A method for quantizing high-dimensional embeddings into discrete tokens.
- Continued Pre-training: Adapting a pre-trained LLM to a specific domain by training it on domain-specific data.
- Generative Retrieval: Using an LLM to generate recommendations directly, rather than simply ranking existing candidates.
Semantic IDs: Addressing Cold-Start and Sparsity
- Challenge: Hash-based item IDs in recommendation systems don't encode item content, leading to cold-start problems and sparsity issues. New items require relearning, and tail items lack sufficient interaction data.
- Solution: Semantic IDs that incorporate multimodal content (visual, text, audio) to represent items.
- Example: Kuaishou's Trainable Multimodal Semantic IDs:
- Model Architecture: A two-tower network with user and item embedding layers.
- Content Input: Uses ResNet for visual encoding, BERT for video descriptions, and VGGish for audio encoding.
- Cluster IDs: Content embeddings are concatenated, and k-means clustering is used to learn cluster IDs (e.g., 1000 clusters for 100 million videos).
- Training: The model learns to map content space (via cluster IDs) to behavioral space.
- Results: Semantic IDs outperform hash-based IDs in clicks and likes, increasing co-start coverage by 3.6% and co-start velocity.
- Benefits: Semantic IDs address cold-start issues and enable recommendations to understand content.
Data Augmentation with LLMs: Enhancing Data Quality and Scale
- Challenge: Obtaining high-quality data at scale is costly and requires significant effort, especially for search and recommendation systems.
- Solution: Using LLMs for synthetic data generation and labeling.
- Example 1: Indeed's Job Recommendation Filtering:
- Problem: Poor job recommendations leading to user unsubscribes. Sparse explicit negative feedback and imprecise implicit feedback.
- Solution: A lightweight classifier to filter bad recommendations.
- Process:
- Experts labeled job recommendations and user pairs.
- Prompted open-source LLMs (Mistral, Llama 2), but performance was poor.
- GPT-4 achieved 90% precision and recall but was too costly and slow (22 seconds).
- GPT-3.5 had poor precision (63%).
- Fine-tuned GPT-3.5, achieving desired precision but still too slow (6.7 seconds).
- Distilled a lightweight classifier on the fine-tuned GPT-3.5 labels, achieving 0.86 AUROC and real-time filtering.
- Outcome: Reduced bad recommendations by 20%, increased application rate by 4%, and decreased unsubscribe rate by 5%.
- Example 2: Spotify's Query Recommendation System:
- Problem: Introducing podcasts and audiobooks to a user base primarily searching for songs and artists.
- Solution: A query recommendation system to generate new queries.
- Process:
- Extracted queries from catalog titles, playlist titles, and search logs.
- Used LLMs to generate natural language queries.
- Ranked new queries alongside immediate search results.
- Outcome: A 9% increase in exploratory queries, driving users to explore new product categories.
- Benefits: LLM-augmented synthetic data provides richer, higher-quality data at scale, especially for tail queries and items, at a lower cost and effort than human annotation.
Unified Models: Streamlining Recommendation Systems
- Challenge: Separate systems for ads, recommendations, and search, with multiple models for different recommendation scenarios, leading to duplicative engineering pipelines and maintenance costs.
- Solution: Unified models that consolidate multiple tasks into a single model.
- Example 1: Netflix's Unified Contextual Ranker (Unicorn):
- Problem: Bespoke models for search, similar item recommendations, and pre-query recommendations, resulting in high operational costs.
- Solution: A unified ranker that takes unified input (user ID, item ID, search query, country, task).
- Process:
- A user foundation model takes in user watch history.
- A context and relevance model takes in the context of videos watched.
- Imputation of missing items (e.g., using the title of the current item as a search query for item-to-item recommendations).
- Outcome: The unified model matched or exceeded the metrics of specialized models on multiple tasks.
- Example 2: Etsy's Unified Embeddings:
- Problem: Helping users get better results from specific or broad queries, given Etsy's constantly changing inventory.
- Solution: Unified embedding and retrieval.
- Model Architecture:
- Product Tower: T5 models for text embeddings (item descriptions) and query embeddings (query-product logs).
- Query Tower: Shared encoders for text tokens, product category tokens, and user location.
- User Preferences: Encoded via query user-scale effect features.
- System Architecture:
- Product and query encoders with a quality vector concatenated to the product embedding.
- A constant vector is added to the query embedding to match the dimension of the product embedding.
- Results: A 2.6% increase in conversion across the entire site and a 5% increase in search purchases.
- Benefits: Unified models simplify systems, improve multiple use cases, and reduce tech debt.
Pinterest's LLM Integration for Search Relevance
- Model Architecture: Cross-encoder structure where query and pin text are concatenated and passed into an LLM for embedding. The embedding is then fed into an MLP layer to produce a five-dimensional vector corresponding to five relevance levels.
- Key Learnings:
- LLMs are good at relevance prediction, with larger models showing substantial improvements over traditional methods.
- Vision language model-generated captions and user actions are useful content annotations.
- LLMs can be used to generate high-quality negative examples for training.
- Model efficiency can be improved through specification, smaller models, and quantization.
- Model Efficiency:
- Specification: Reduce the number of heads in transformers, reduce the number of MLPs.
- Smaller Models: Distill larger models into smaller models step by step.
- Quantization: Use mixed precision, with the LM head in FP32.
- Results:
- Significant reduction in latency (7x) and increase in throughput (30x).
Instacart's Use of LLMs for Search and Discovery
- Challenges:
- Overly broad queries and very specific queries lack sufficient engagement data.
- Supporting new item discovery.
- Solutions:
- LLMs to uplevel query understanding.
- Hybrid approach combining LLMs with traditional models.
- Query Understanding:
- Query to Category Classifier:
- LLMs predict relevant categories for a query.
- Top converting categories are fed as additional context to the LLM.
- Precision improved by 18 percentage points and recall improved by 70 percentage points for tail queries.
- Query Rewrites Model:
- LLMs generate precise rewrites (substitute, broad, synonymous).
- Large drop in the number of queries without any results.
- Query to Category Classifier:
- Discovery-Oriented Content:
- LLMs generate substitute and complementary results.
- Augment prompts with Instacart domain knowledge (top converting categories, query understanding annotations, subsequent queries).
- Improved engagement and revenue.
- Serving:
- Precompute outputs for head and torso queries and cache the data.
- Fall back to existing models for the long tail of queries.
- Key Takeaways:
- LLM's world knowledge is important for improving query understanding predictions.
- Combining domain knowledge with LLMs is crucial for success.
- Evaluating content and query predictions is important and difficult.
YouTube's Adaptation of Gemini for Recommendations
- Problem: Building a recommendation system that can handle the scale and freshness requirements of YouTube.
- Solution: Adapting a Gemini checkpoint for YouTube recommendations.
- Process:
- Tokenize Videos:
- Extract features from videos (title, description, transcript, audio, video frames).
- Create a multi-dimensional embedding.
- Quantize the embedding using RQVE to create semantic IDs.
- Adapt the LLM:
- Link text and semantic IDs.
- Train the model to understand sequences of watches.
- Prompt with User Information:
- Construct personalized prompts with user demographics, activity, and actions.
- Train task-specific models.
- Tokenize Videos:
- Generative Retrieval:
- Construct a prompt for each user and have the model decode video recommendations as semantic IDs.
- Unique recommendations, especially for users with limited data.
- Challenges:
- Vocabulary and size of the corpus (billions of videos).
- Freshness of videos.
- Scale (billions of daily active users).
- LLM and Rexus Recipe:
- Tokenize content.
- Adapt the LLM.
- Prompt with user information.
Synthesis/Conclusion
The talks highlight the transformative potential of LLMs in recommendation systems. Semantic IDs address cold-start and sparsity issues by encoding item content. Data augmentation with LLMs enriches datasets, improving model performance. Unified models streamline systems by consolidating multiple tasks into a single model. The key to success lies in combining LLMs with domain-specific knowledge and carefully addressing challenges related to scale, freshness, and evaluation. The future of recommendation systems involves more interactive experiences, personalized content generation, and a blurring of lines between search and recommendation.
AI summaries can miss context or contain errors. Check important details against the original video.