Analysis code for "When should recommender systems explore? Dynamic user modeling from randomized exposure logs"
收藏资源简介:
Reproducible analysis notebook accompanying the manuscript "When should recommender systems explore? Dynamic user modeling from randomized exposure logs" (submitted to User Modeling and User-Adapted Interaction). A single Jupyter notebook regenerates every table, figure and number in the paper from the public KuaiRand-Pure dataset (Gao et al., 2022; Zenodo record 10439422). The data are not redistributed here. The notebook implements: leakage-safe features computed from each user's observed candidate-pool history (with automated leakage tests), a temporal train/validation/test split, static and dynamic engagement models (logistic regression, LightGBM, GRU benchmark) with isotonic calibration, user-level cluster bootstrap inference, an exploration-receptivity score (ERS), cross-fitted doubly robust (AIPW) estimation of the effect of off-profile exposure using the randomized exposures, off-policy evaluation of exploration policies (IPS, SNIPS, DR), ablations, subgroup analyses and SHAP interpretation. Runs on Kaggle (CPU, about 1 hour in FULL mode) or locally; see README for instructions. Set RUN_MODE = "FULL" to reproduce the paper (the default "FAST" mode is a smoke test on a user sample). The export step writes a SHA-256 manifest of all inputs and outputs.



