Teo Narboneta Zosa, Chingis Oinar – Robust AI Search Ranking for Radical C2C Marketplace Growth
By Plain Schwarz
Key Concepts:
- E-commerce search ranking
- Learning to Rank (LTR)
- Implicit judgments
- Offline evaluation (eval)
- NDCG (Normalized Discounted Cumulative Gain)
- Purchase NDCG
- High signal filtering
- Data debiasing (position bias)
- Intervention harvesting
- Model building (pointwise, pairwise, listwise approaches)
- Context features, document features
- Feature transformation (Japanese tokenization, z-score normalization, L1/LP transformation)
- Model architecture (MLP)
- A/B testing
- User intent
- Search session
1. Search at Mercari:
- Mercari's vision is to create a global marketplace where anyone can buy and sell easily.
- The platform has over 23 million monthly active users in Japan.
- There are over 3 billion item listings, with $6.5 billion in total sales in the last fiscal year.
- Mercari's UI emphasizes simplicity, showing pictures, prices, and a "sold" marker.
- Search is critical, handling tens of thousands of QPS with low hundreds of milliseconds P99 latency.
- The search system has two stages: retrieval (Elasticsearch) and post-retrieval ML ranking.
- The focus of the talk is on the ML ranking component.
- LTR is used to train an ML model to learn a ranking function.
- The goal of the ML ranking is to order items based on user intent, not just query relevance.
- The aim is to remove the need for users to become SEO experts.
2. Data Set Construction and Offline Evaluation:
- The key to building an effective system is good evals and data literacy.
- Data set construction starts with implicit judgments derived from user search logs.
- Relevance labels are ordered from most to least relevant based on user actions: purchases, adds to cart, comments, likes, and clicks.
- Initially, 5 days of training data and 1 day of validation data are used to prevent temporal leakage.
- Standard search quality metrics like NDCG are used for offline evaluation.
- Qualitative spot checks ("Vibe checks") are performed to ensure rankings look reasonable.
- Weights and Biases is used for the evaluation dashboard.
- Clicks were found to be a noisy signal, over-represented in the data set.
- Removing clicks improved correlation between offline and online metrics.
- Custom metrics, especially "purchase NDCG," are crucial for delivering business value.
- Purchase NDCG has a 100% hit rate on predicting statistically significant A/B test results.
- Model selection is based on purchase NDCG, while training optimizes for NDCG over relevance labels.
3. High Signal Filtering:
- Purchase NDCG allows filtering to serps with only purchase labels, improving model performance.
- Training over a whole month of data with seven days for validation addresses time seasonality effects.
- Sparse labels are addressed by including other user purchase, checkout, and completed purchase labels.
- Clicks on sold items are added as a label to inspire sellers and improve purchase NDCG.
- The model was previously ranking sold items at the bottom, but this change aims to improve seller-side metrics.
4. Model Building:
- Three main approaches to LTR: pointwise, pairwise, and listwise.
- Features are categorized into context features (query, Ser statistics) and document features (price, title, recency).
- Japanese-oriented tokenization methods (e.g., MeCab) are used.
- Z-score normalization is applied to Ser statistics.
- L1/LP transformation is used for numerical features.
- The model architecture is a basic MLP model.
5. Debiasing:
- Elasticsearch was initially used as the ranking function.
- Gaussian noise is added to the item rank feature to regularize the model.
- Position bias is addressed using intervention harvesting.
- Data is collected from A/B tests to find documents that were shuffled.
- Click-through rate is computed at each position and normalized by the expected click-through rate.
- Weights are computed to represent the bias for each rank.
- TensorFlow Ranking is used to apply document weights (estimated position bias).
- Applying weights to labels instead of the loss function yields better results.
6. Takeaways and What's Next:
- Foundations for getting models into production are essential.
- Tight evaluation and data enrichment feedback loops are crucial.
- Start simple and iterate quickly.
- The AI features released have had a significant impact on the company's growth.
- Future work includes using LLMs as quasi-explicit raters and addressing spammy items.
- Focus on applying proven use cases and methods to other projects.
7. Notable Quotes:
- "The key to building an effective and robust system is good evals and data literacy."
- "Don't disdain the small things, really just start somewhere."
8. Technical Terms:
- QPS: Queries Per Second
- P99 Latency: The 99th percentile of latency, meaning 99% of requests are served within that time.
- Temporal Leakage: When data from the future is inadvertently used to train a model, leading to over-optimistic results.
- Serp: Search Engine Results Page
- CTR: Click-Through Rate
9. Logical Connections:
- The talk progresses from the high-level overview of Mercari's search system to the specifics of data set construction, model building, and debiasing.
- The importance of offline evaluation is emphasized as a prerequisite for successful A/B testing.
- The challenges of noisy and sparse data are addressed with specific techniques like high signal filtering and the inclusion of additional labels.
- The discussion of debiasing builds upon the understanding of position bias and the limitations of naive approaches like random shuffling.
10. Synthesis/Conclusion:
The presentation details Mercari's approach to building a robust e-commerce search ranking system. Key takeaways include the importance of data quality, rigorous evaluation, and addressing biases in user behavior. The journey from a basic Elasticsearch-based system to an ML-powered ranking system with custom metrics and debiasing techniques is highlighted. The emphasis on starting simple, iterating quickly, and focusing on business value is a recurring theme. The presentation concludes with a look at future directions, including the use of LLMs and the exploration of new user experience metrics.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development