Teo Narboneta Zosa, Chingis Oinar – Robust AI Search Ranking for Radical C2C Marketplace Growth

By Plain Schwarz

Share:

Key Concepts:

  • E-commerce search ranking
  • Learning to Rank (LTR)
  • Implicit judgments
  • Offline evaluation (eval)
  • NDCG (Normalized Discounted Cumulative Gain)
  • Purchase NDCG
  • High signal filtering
  • Data debiasing (position bias)
  • Intervention harvesting
  • Model building (pointwise, pairwise, listwise approaches)
  • Context features, document features
  • Feature transformation (Japanese tokenization, z-score normalization, L1/LP transformation)
  • Model architecture (MLP)
  • A/B testing
  • User intent
  • Search session

1. Search at Mercari:

  • Mercari's vision is to create a global marketplace where anyone can buy and sell easily.
  • The platform has over 23 million monthly active users in Japan.
  • There are over 3 billion item listings, with $6.5 billion in total sales in the last fiscal year.
  • Mercari's UI emphasizes simplicity, showing pictures, prices, and a "sold" marker.
  • Search is critical, handling tens of thousands of QPS with low hundreds of milliseconds P99 latency.
  • The search system has two stages: retrieval (Elasticsearch) and post-retrieval ML ranking.
  • The focus of the talk is on the ML ranking component.
  • LTR is used to train an ML model to learn a ranking function.
  • The goal of the ML ranking is to order items based on user intent, not just query relevance.
  • The aim is to remove the need for users to become SEO experts.

2. Data Set Construction and Offline Evaluation:

  • The key to building an effective system is good evals and data literacy.
  • Data set construction starts with implicit judgments derived from user search logs.
  • Relevance labels are ordered from most to least relevant based on user actions: purchases, adds to cart, comments, likes, and clicks.
  • Initially, 5 days of training data and 1 day of validation data are used to prevent temporal leakage.
  • Standard search quality metrics like NDCG are used for offline evaluation.
  • Qualitative spot checks ("Vibe checks") are performed to ensure rankings look reasonable.
  • Weights and Biases is used for the evaluation dashboard.
  • Clicks were found to be a noisy signal, over-represented in the data set.
  • Removing clicks improved correlation between offline and online metrics.
  • Custom metrics, especially "purchase NDCG," are crucial for delivering business value.
  • Purchase NDCG has a 100% hit rate on predicting statistically significant A/B test results.
  • Model selection is based on purchase NDCG, while training optimizes for NDCG over relevance labels.

3. High Signal Filtering:

  • Purchase NDCG allows filtering to serps with only purchase labels, improving model performance.
  • Training over a whole month of data with seven days for validation addresses time seasonality effects.
  • Sparse labels are addressed by including other user purchase, checkout, and completed purchase labels.
  • Clicks on sold items are added as a label to inspire sellers and improve purchase NDCG.
  • The model was previously ranking sold items at the bottom, but this change aims to improve seller-side metrics.

4. Model Building:

  • Three main approaches to LTR: pointwise, pairwise, and listwise.
  • Features are categorized into context features (query, Ser statistics) and document features (price, title, recency).
  • Japanese-oriented tokenization methods (e.g., MeCab) are used.
  • Z-score normalization is applied to Ser statistics.
  • L1/LP transformation is used for numerical features.
  • The model architecture is a basic MLP model.

5. Debiasing:

  • Elasticsearch was initially used as the ranking function.
  • Gaussian noise is added to the item rank feature to regularize the model.
  • Position bias is addressed using intervention harvesting.
  • Data is collected from A/B tests to find documents that were shuffled.
  • Click-through rate is computed at each position and normalized by the expected click-through rate.
  • Weights are computed to represent the bias for each rank.
  • TensorFlow Ranking is used to apply document weights (estimated position bias).
  • Applying weights to labels instead of the loss function yields better results.

6. Takeaways and What's Next:

  • Foundations for getting models into production are essential.
  • Tight evaluation and data enrichment feedback loops are crucial.
  • Start simple and iterate quickly.
  • The AI features released have had a significant impact on the company's growth.
  • Future work includes using LLMs as quasi-explicit raters and addressing spammy items.
  • Focus on applying proven use cases and methods to other projects.

7. Notable Quotes:

  • "The key to building an effective and robust system is good evals and data literacy."
  • "Don't disdain the small things, really just start somewhere."

8. Technical Terms:

  • QPS: Queries Per Second
  • P99 Latency: The 99th percentile of latency, meaning 99% of requests are served within that time.
  • Temporal Leakage: When data from the future is inadvertently used to train a model, leading to over-optimistic results.
  • Serp: Search Engine Results Page
  • CTR: Click-Through Rate

9. Logical Connections:

  • The talk progresses from the high-level overview of Mercari's search system to the specifics of data set construction, model building, and debiasing.
  • The importance of offline evaluation is emphasized as a prerequisite for successful A/B testing.
  • The challenges of noisy and sparse data are addressed with specific techniques like high signal filtering and the inclusion of additional labels.
  • The discussion of debiasing builds upon the understanding of position bias and the limitations of naive approaches like random shuffling.

10. Synthesis/Conclusion:

The presentation details Mercari's approach to building a robust e-commerce search ranking system. Key takeaways include the importance of data quality, rigorous evaluation, and addressing biases in user behavior. The journey from a basic Elasticsearch-based system to an ML-powered ranking system with custom metrics and debiasing techniques is highlighted. The emphasis on starting simple, iterating quickly, and focusing on business value is a recurring theme. The presentation concludes with a look at future directions, including the use of LLMs and the exploration of new user experience metrics.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video