AI Engineer World’s Fair 2025 - LLM RECSYS

AI EngineerAbout 8 min readJun 8, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Semantic IDs: Content-aware item identifiers that encode item characteristics, addressing cold-start and sparsity issues.
  • Data Augmentation: Using Language Models (LLMs) to generate synthetic data and labels, enriching datasets for search and recommendation systems.
  • Unified Models: Consolidating multiple recommendation, search, and advertising models into a single, versatile model to reduce redundancy and improve efficiency.
  • Co-start: The problem of recommending new or unpopular items with limited interaction data.
  • Co-star Velocity: The rate at which new items reach a certain threshold of views or interactions.
  • Trainable Multimodal Semantic IDs: A method of combining static content embeddings (visual, text, audio) with dynamic user behavior to create semantic IDs.
  • Exploratory Search: Helping users discover new categories or items beyond their usual searches.
  • Query Recommendation System: Recommending related search queries to users to facilitate exploratory search.
  • Multi-task Learning: Training a single model to perform multiple related tasks simultaneously.
  • Unified Contextual Ranker: A unified model that ranks items for search, similar item recommendations, and pre-query recommendations.
  • Quality Vector: A vector representing the quality of an item based on ratings, freshness, and conversion rate.
  • RQVE (Residual Quantized Vector Embedding): A method for quantizing high-dimensional embeddings into discrete tokens.
  • Continued Pre-training: Adapting a pre-trained LLM to a specific domain by training it on domain-specific data.
  • Generative Retrieval: Using an LLM to generate recommendations directly, rather than simply ranking existing candidates.

Semantic IDs: Addressing Cold-Start and Sparsity

  • Challenge: Hash-based item IDs in recommendation systems don't encode item content, leading to cold-start problems and sparsity issues. New items require relearning, and tail items lack sufficient interaction data.
  • Solution: Semantic IDs that incorporate multimodal content (visual, text, audio) to represent items.
  • Example: Kuaishou's Trainable Multimodal Semantic IDs:
    • Model Architecture: A two-tower network with user and item embedding layers.
    • Content Input: Uses ResNet for visual encoding, BERT for video descriptions, and VGGish for audio encoding.
    • Cluster IDs: Content embeddings are concatenated, and k-means clustering is used to learn cluster IDs (e.g., 1000 clusters for 100 million videos).
    • Training: The model learns to map content space (via cluster IDs) to behavioral space.
    • Results: Semantic IDs outperform hash-based IDs in clicks and likes, increasing co-start coverage by 3.6% and co-start velocity.
  • Benefits: Semantic IDs address cold-start issues and enable recommendations to understand content.

Data Augmentation with LLMs: Enhancing Data Quality and Scale

  • Challenge: Obtaining high-quality data at scale is costly and requires significant effort, especially for search and recommendation systems.
  • Solution: Using LLMs for synthetic data generation and labeling.
  • Example 1: Indeed's Job Recommendation Filtering:
    • Problem: Poor job recommendations leading to user unsubscribes. Sparse explicit negative feedback and imprecise implicit feedback.
    • Solution: A lightweight classifier to filter bad recommendations.
    • Process:
      1. Experts labeled job recommendations and user pairs.
      2. Prompted open-source LLMs (Mistral, Llama 2), but performance was poor.
      3. GPT-4 achieved 90% precision and recall but was too costly and slow (22 seconds).
      4. GPT-3.5 had poor precision (63%).
      5. Fine-tuned GPT-3.5, achieving desired precision but still too slow (6.7 seconds).
      6. Distilled a lightweight classifier on the fine-tuned GPT-3.5 labels, achieving 0.86 AUROC and real-time filtering.
    • Outcome: Reduced bad recommendations by 20%, increased application rate by 4%, and decreased unsubscribe rate by 5%.
  • Example 2: Spotify's Query Recommendation System:
    • Problem: Introducing podcasts and audiobooks to a user base primarily searching for songs and artists.
    • Solution: A query recommendation system to generate new queries.
    • Process:
      1. Extracted queries from catalog titles, playlist titles, and search logs.
      2. Used LLMs to generate natural language queries.
      3. Ranked new queries alongside immediate search results.
    • Outcome: A 9% increase in exploratory queries, driving users to explore new product categories.
  • Benefits: LLM-augmented synthetic data provides richer, higher-quality data at scale, especially for tail queries and items, at a lower cost and effort than human annotation.

Unified Models: Streamlining Recommendation Systems

  • Challenge: Separate systems for ads, recommendations, and search, with multiple models for different recommendation scenarios, leading to duplicative engineering pipelines and maintenance costs.
  • Solution: Unified models that consolidate multiple tasks into a single model.
  • Example 1: Netflix's Unified Contextual Ranker (Unicorn):
    • Problem: Bespoke models for search, similar item recommendations, and pre-query recommendations, resulting in high operational costs.
    • Solution: A unified ranker that takes unified input (user ID, item ID, search query, country, task).
    • Process:
      1. A user foundation model takes in user watch history.
      2. A context and relevance model takes in the context of videos watched.
      3. Imputation of missing items (e.g., using the title of the current item as a search query for item-to-item recommendations).
    • Outcome: The unified model matched or exceeded the metrics of specialized models on multiple tasks.
  • Example 2: Etsy's Unified Embeddings:
    • Problem: Helping users get better results from specific or broad queries, given Etsy's constantly changing inventory.
    • Solution: Unified embedding and retrieval.
    • Model Architecture:
      • Product Tower: T5 models for text embeddings (item descriptions) and query embeddings (query-product logs).
      • Query Tower: Shared encoders for text tokens, product category tokens, and user location.
      • User Preferences: Encoded via query user-scale effect features.
    • System Architecture:
      • Product and query encoders with a quality vector concatenated to the product embedding.
      • A constant vector is added to the query embedding to match the dimension of the product embedding.
    • Results: A 2.6% increase in conversion across the entire site and a 5% increase in search purchases.
  • Benefits: Unified models simplify systems, improve multiple use cases, and reduce tech debt.

Pinterest's LLM Integration for Search Relevance

  • Model Architecture: Cross-encoder structure where query and pin text are concatenated and passed into an LLM for embedding. The embedding is then fed into an MLP layer to produce a five-dimensional vector corresponding to five relevance levels.
  • Key Learnings:
    • LLMs are good at relevance prediction, with larger models showing substantial improvements over traditional methods.
    • Vision language model-generated captions and user actions are useful content annotations.
    • LLMs can be used to generate high-quality negative examples for training.
    • Model efficiency can be improved through specification, smaller models, and quantization.
  • Model Efficiency:
    • Specification: Reduce the number of heads in transformers, reduce the number of MLPs.
    • Smaller Models: Distill larger models into smaller models step by step.
    • Quantization: Use mixed precision, with the LM head in FP32.
  • Results:
    • Significant reduction in latency (7x) and increase in throughput (30x).

Instacart's Use of LLMs for Search and Discovery

  • Challenges:
    • Overly broad queries and very specific queries lack sufficient engagement data.
    • Supporting new item discovery.
  • Solutions:
    • LLMs to uplevel query understanding.
    • Hybrid approach combining LLMs with traditional models.
  • Query Understanding:
    • Query to Category Classifier:
      • LLMs predict relevant categories for a query.
      • Top converting categories are fed as additional context to the LLM.
      • Precision improved by 18 percentage points and recall improved by 70 percentage points for tail queries.
    • Query Rewrites Model:
      • LLMs generate precise rewrites (substitute, broad, synonymous).
      • Large drop in the number of queries without any results.
  • Discovery-Oriented Content:
    • LLMs generate substitute and complementary results.
    • Augment prompts with Instacart domain knowledge (top converting categories, query understanding annotations, subsequent queries).
    • Improved engagement and revenue.
  • Serving:
    • Precompute outputs for head and torso queries and cache the data.
    • Fall back to existing models for the long tail of queries.
  • Key Takeaways:
    • LLM's world knowledge is important for improving query understanding predictions.
    • Combining domain knowledge with LLMs is crucial for success.
    • Evaluating content and query predictions is important and difficult.

YouTube's Adaptation of Gemini for Recommendations

  • Problem: Building a recommendation system that can handle the scale and freshness requirements of YouTube.
  • Solution: Adapting a Gemini checkpoint for YouTube recommendations.
  • Process:
    1. Tokenize Videos:
      • Extract features from videos (title, description, transcript, audio, video frames).
      • Create a multi-dimensional embedding.
      • Quantize the embedding using RQVE to create semantic IDs.
    2. Adapt the LLM:
      • Link text and semantic IDs.
      • Train the model to understand sequences of watches.
    3. Prompt with User Information:
      • Construct personalized prompts with user demographics, activity, and actions.
      • Train task-specific models.
  • Generative Retrieval:
    • Construct a prompt for each user and have the model decode video recommendations as semantic IDs.
    • Unique recommendations, especially for users with limited data.
  • Challenges:
    • Vocabulary and size of the corpus (billions of videos).
    • Freshness of videos.
    • Scale (billions of daily active users).
  • LLM and Rexus Recipe:
    1. Tokenize content.
    2. Adapt the LLM.
    3. Prompt with user information.

Synthesis/Conclusion

The talks highlight the transformative potential of LLMs in recommendation systems. Semantic IDs address cold-start and sparsity issues by encoding item content. Data augmentation with LLMs enriches datasets, improving model performance. Unified models streamline systems by consolidating multiple tasks into a single model. The key to success lies in combining LLMs with domain-specific knowledge and carefully addressing challenges related to scale, freshness, and evaluation. The future of recommendation systems involves more interactive experiences, personalized content generation, and a blurring of lines between search and recommendation.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.