What We Learned from Using LLMs in Pinterest — Mukuntha Narayanan, Han Wang, Pinterest

AI EngineerAbout 4 min readJul 17, 2025Watch original
THE SUMMARYAI-generated

Pinterest Search Relevance with Large Language Models: A Summary

Key Concepts:

  • Semantic Relevance Modeling: Using machine learning to determine how relevant a pin is to a user's search query.
  • Large Language Models (LLMs): Powerful neural networks trained on vast amounts of text data, used for relevance prediction and content representation.
  • Cross-Encoder vs. Bi-Encoder: Two different model architectures for processing query and pin information. Cross-encoders process the query and pin together, while bi-encoders process them separately.
  • Knowledge Distillation: Training a smaller "student" model to mimic the behavior of a larger, more complex "teacher" model (LLM).
  • Semi-Supervised Learning: Training a model using a combination of labeled data (human annotations) and unlabeled data (search logs labeled by the teacher model).
  • Visual Language Model (VLM) Captions: Automatically generated text descriptions of images, used as features for pin representation.
  • User Action Based Features: Data derived from user interactions with pins, such as board titles and search queries that led to engagement.
  • Re-ranking: The stage in the search pipeline where the initial set of candidate pins is reordered based on relevance scores.
  • Embedding: A numerical representation of text or images that captures their semantic meaning.
  • BM25: A ranking function used in information retrieval to estimate the relevance of documents to a given search query.

1. LLMs for Relevance Prediction

  • Main Point: LLMs are highly effective at predicting the relevance of pins to search queries.
  • Model Architecture: A cross-encoder architecture is used, where the query and pin text are concatenated and fed into an LLM to generate an embedding. This embedding is then passed through an MLP layer to predict a five-dimensional vector representing the five relevance levels.
  • Training: Open-source LLMs are fine-tuned using Pinterest's internal data to adapt them to Pinterest content.
  • Results: LLMs significantly outperform traditional methods like Search Sage (Pinterest's in-house embedding model).
    • An 8 billion parameter Llama mastery model achieved a 12% improvement over a multilingual BERT-based model and a 20% improvement over Search Sage.
  • Key Takeaway: Using larger and more advanced LLMs leads to better relevance prediction performance.

2. Content Annotations: VLM Captions and User Actions

  • Main Point: VLM-generated captions and user action-based features are valuable sources of content annotations that improve relevance prediction.
  • Features Used:
    • Pin title and description
    • VLM-generated synthetic image captions
    • Board titles where the pin has been saved
    • Queries that led to high engagement with the pin on search
  • Ablation Study: Removing each feature type sequentially to assess its impact on performance.
    • VLM captions provide a solid baseline.
    • Adding more text features leads to performance improvements.
    • User action-based features (board titles, queries) are particularly useful.
  • Key Takeaway: Enriching pin representations with diverse features, especially those derived from user behavior, enhances the model's understanding of the content.

3. Knowledge Distillation for Productionization

  • Main Point: Knowledge distillation is used to create a smaller, more efficient student model that can be deployed in production without incurring excessive computational costs.
  • Process:
    1. Teacher Model (LLM): Trained on a small set of human-labeled data.
    2. Data Generation: Daily search logs are sampled and labeled by the teacher model, creating a large dataset of pseudo-labeled data.
    3. Student Model: Trained on the pseudo-labeled data to predict the same five-scale relevance scores as the teacher model.
  • Student Model Architecture: A bi-encoder architecture is used for the student model, which allows for offline inference and caching of pin embeddings.
    • Pin embeddings are pre-computed offline and updated only when the pin content changes significantly.
    • Query embeddings are computed in real-time online.
  • Features in Student Model:
    • LLM embeddings (bi-encoder)
    • Graph Sage embeddings
    • BM25 scores
  • Scaling: Offline inference and caching of pin embeddings enable the model to scale to billions of pins. Query embedding caching achieves an 85% hit rate.
  • Results: The student model achieves relevance gains internationally, even though the teacher model was primarily trained on US data. Search fulfillment (user engagement) also increases.

4. Semantic Representations from Relevance Tuning

  • Main Point: Relevance-tuned LLMs produce rich semantic representations that can be used for various tasks beyond search.
  • Application: The pin and query embeddings from the production relevance student model are used as general-purpose content representations across Pinterest.
  • Benefits: These embeddings capture semantic meaning effectively and improve performance in related pins, home feed, and other surfaces.

5. Conclusion

The presentation highlights the successful integration of LLMs into Pinterest search to improve relevance and user experience. Key takeaways include the effectiveness of LLMs for relevance prediction, the importance of rich content annotations (VLM captions and user actions), the use of knowledge distillation for efficient productionization, and the value of relevance-tuned embeddings for general-purpose content representation. The use of semi-supervised learning allows Pinterest to leverage vast amounts of unlabeled data to train high-performing models that scale globally.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.