THE SUMMARYAI-generated
Instacart Search and Discovery with LLMs: A Summary
Key Concepts:
- Query Understanding: The process of interpreting user search queries to improve retrieval and ranking.
- Long Tail Queries: Infrequent, specific search queries with limited engagement data.
- Cold Start Problem: Difficulty in ranking new or less popular products due to lack of engagement data.
- LLMs (Large Language Models): Powerful AI models used for natural language processing and generation.
- Domain Knowledge: Specific knowledge related to Instacart's products, user behavior, and business goals.
- Hybrid Approach: Combining LLMs with existing models and techniques.
- Discovery-Oriented Content: Showing related or alternative products to encourage exploration and larger basket sizes.
- Query Rewrites: Modifying user queries to improve search results, especially for retailers with smaller catalogs.
- Prompt Engineering: Designing effective prompts for LLMs to generate desired outputs.
- Batch Mode vs. Real-time: Processing data offline in batches versus generating results in real-time.
1. Importance of Search in Grocery E-commerce
- Search is crucial for grocery e-commerce, as customers often have long shopping lists with both restocking purchases and new items.
- Search needs to efficiently find specific products and enable new product discovery.
- New product discovery benefits customers, advertisers (showcasing new products), and the platform (larger basket sizes).
2. Challenges with Conventional Search Engines
- Overly Broad Queries: Queries like "snacks" return too many products, making it difficult to collect engagement data for ranking.
- Very Specific Queries: Queries like "unsweetened plant-based yogurt" are infrequent, resulting in insufficient engagement data.
- New Item Discovery: Customers want a similar experience to browsing a physical store aisle, but finding related products through search is often difficult.
- Precision vs. Recall: Improving recall (finding relevant products) often comes at the cost of precision (showing only relevant products).
3. Using LLMs to Enhance Query Understanding
- Query Understanding Module: The initial stage of the search stack, requiring accurate outputs for retrieval and ranking.
- Models within the Module: Query normalization, query tagging, query classification, category classification, etc.
- Query to Category Classifier: Maps a query to relevant categories in a taxonomy (10,000 labels, 6,000 commonly used). This is a multilabel classification problem.
- Previous Models: FastText-based neural network and NPMI model (statistical co-occurrence).
- Limitations: Low coverage for tail queries due to lack of engagement data. Bird-based models showed limited improvement for the increased latency.
- LLM Approach: Initially, feeding queries and the taxonomy to the LLM yielded decent results but poor online AB test performance.
- Example: "Protein" query was interpreted as protein foods (chicken, tofu) instead of protein shakes/bars.
- Improved LLM Approach: Providing top converting categories for each query as additional context to the LLM.
- Example: "Wner soda" was correctly identified as a brand of ginger ale instead of just a fruit-flavored soda.
- Results: Significant improvement in precision (18 percentage points) and recall (70 percentage points) for tail queries.
- Prompt: Simple prompt passing in top converted categories as context with guidelines for the LLM.
- Query Rewrites Model: Rewrites queries to improve search results, especially when retailers have limited catalogs.
- Example: Rewriting "1% milk" to "milk" to ensure results are returned.
- Previous Approach: Trained on engagement data, performed poorly on tail queries.
- LLM Approach: Generating precise rewrites (substitute, broad, synonymous).
- Example: For "avocado oil," the LLM generated "olive oil" (substitute), "healthy cooking oil" (broader), and "avocado extract" (synonymous).
- Results: Large drop in the number of queries with no results.
- Offline Improvements: Moving from simpler models to better models improved human evaluation scores.
4. Scoring and Serving LLM-Generated Data
- Instacart's Query Pattern: Fat head and torso with a long tail of queries.
- Hybrid Approach:
- Pre-computing outputs for head and torso queries offline and caching the data.
- Serving from the cache online with low latency impact.
- Falling back to existing models for the long tail of queries.
- Future Plans: Replacing existing models for the long tail with a distilled Llama model.
5. Consolidating Models and Leveraging Context
- Managing multiple models in the query understanding module is complex.
- Consolidating models into a single SLM (or LLM) can improve consistency.
- Example: The query "hmm" was incorrectly corrected to "hummus" by the spell corrector, leading to confusing results. A unified model could avoid this.
- LLMs can leverage extra context to understand the customer's mission (e.g., buying ingredients for a recipe) and generate relevant content.
6. Using LLMs for Discovery-Oriented Content
- Problem: Users couldn't easily find related products after adding an item to their cart.
- Solution: Using LLMs to generate substitute results (when no exact results are found) and complementary results (at the bottom of the search results page).
- Example: For "swordfish," showing seafood alternatives like tilapia. For "sushi," showing Asian cooking ingredients or Japanese drinks.
- Requirements:
- Content must be incremental to existing solutions (no duplicates).
- Content must be aligned with Instacart's domain knowledge (e.g., "dishes" refers to cookware, not food).
- Initial Approach: Basic generation using a prompt asking the LLM to generate complementary and substitute items.
- Limitation: LLM answers were common sense but not aligned with Instacart user behavior.
- Example: "Protein" query resulted in chicken/tofu instead of protein bars/shakes.
- Limitation: LLM answers were common sense but not aligned with Instacart user behavior.
- Improved Approach: Augmenting the prompt with Instacart domain knowledge:
- Top converting categories for the query.
- Annotations from the query understanding model (brand, dietary attribute).
- Subsequent queries that users performed after the initial query.
- Serving:
- Calling the LLM in batch mode for historical search logs.
- Storing query-content metadata and potential products in a feature store.
- Quick lookup from the feature store online.
- Challenges:
- Aligning generation with business metrics (revenue).
- Improving the ranking of the content (diversity-based ranking).
- Evaluating the content (correctness, adherence to Instacart's needs).
7. Key Takeaways
- LLM's world knowledge is crucial for improving query understanding, especially for tail queries.
- Combining Instacart's domain knowledge with LLMs is essential for achieving topline wins.
- Evaluating the content and query predictions is more important and difficult than anticipated. LLM as a judge was used for this.
8. Questions and Answers
- Long Natural Language Queries: Instacart has launched "Ask Instacart" to map natural language queries to search intent.
- Learnings from "Ask Instacart":
- Injecting Instacart context is crucial.
- Robust automated evaluation pipeline is important.
- Passing context from the LLM to downstream systems is essential.
9. Synthesis/Conclusion
Instacart successfully leveraged LLMs to transform its search and discovery experience. By combining the power of LLMs with its own domain knowledge and a hybrid serving approach, Instacart improved query understanding, enabled new product discovery, and ultimately drove business growth. The key to success was not just using LLMs, but carefully engineering prompts, evaluating results, and integrating LLMs into the existing infrastructure.
AI summaries can miss context or contain errors. Check important details against the original video.