Fair-by-Design Retail Media Interaction Dataset
收藏资源简介:
An 18-month corpus of 2.3M user–item interactions from a major e-commerce platform covering 847k users, 156k products, and 15 categories, built for fairness-aware personalization under signal loss; it aggregates privacy-preserving, non-PII signals—product reviews with sentiment and helpfulness, public search-trend indices, social-engagement metrics for brand content, and product/category metadata (e.g., pricing, popularity)—with quality filters (≥5 interactions per user; ≥10 reviews per product) and strong coverage (89.3% sentiment, 76.8% search, 64.2% social). Identifiers are removed or hashed, demographic attributes are inferred via behavioral proxies, and the dataset is designed to evaluate accuracy–fairness trade-offs for CLV-oriented recommender models.



