User Engagement Shift Modeling for Short Format Recommendation System: A Deep Reinforcement Learning Framework Based on Adaptive Time Discounting
收藏资源简介:
# ATD-DRL: Adaptive Time-Discounting Deep Reinforcement Learning Implementation and experimental artefacts for the adaptive time-discountingdeep reinforcement learning framework applied to user engagement shiftmodelling in short-format recommendation systems. The repository hosts two archives that together provide the full experimentalpackage: `atd_drl_code.zip` (the runnable source tree) and`atd_drl_data.zip` (the interaction data, evaluation results and renderedfigures). ## Repository Layout ```atd-drl/|-- atd_drl_code.zip # Source code package (see the code README inside)|-- atd_drl_data.zip # Experimental data package (see the data README inside)|-- README.md # This document``` ## Method Overview The framework raises the discount factor from a scalar hyperparameter to alearnable function conditioned on state embeddings, drift intensity, andconversation phase. Three architectural components work in concert: 1. **Drift-aware dual-stream encoder.** A Transformer over the k-step interaction sequence produces the short-term embedding, and a state-space model over session-level pooled histories produces the long-term embedding. The two streams are combined by a learned gate. The residual of a dynamic prediction head, smoothed by exponential moving average, serves as an unsupervised drift-intensity signal.2. **Adaptive discount network.** A lightweight multi-layer perceptron maps the triple (state, drift, phase) through a sigmoid to a discount value in the open interval (gamma_min, gamma_max). Two regularisers are applied: a smoothing term that prevents adjacent-step jumps and a prior soft constraint that lowers gamma as drift intensity rises.3. **Actor-Critic-gamma cooperative training.** The policy, twin-Q value network, and discount network are updated jointly with IPS-weighted temporal-difference loss, twin-Q suppression, Polyak averaging, and a warm-up phase during which the discount is frozen at a baseline value. ## Datasets The framework is evaluated on three public short-format datasets: | Dataset | Users | Items | Interactions | Avg. session length | Mean drift intensity ||---------|-------|-------|--------------|---------------------|----------------------|| KuaiRec | 1,411 | 3,327 | 4.68 M | 36.4 | 0.183 || MicroVideo-1.7M | 1,028,500 | 1,704,880 | 12.74 M | 47.2 | 0.214 || MIND-Short | 94,057 | 65,238 | 24.16 M | 18.7 | 0.152 | The interaction files used for evaluation are provided in the data package. ## Main Results | Method | KuaiRec NDCG@10 | KuaiRec Retention | MicroVideo NDCG@10 | MicroVideo Retention | MIND NDCG@10 | MIND Retention ||--------|-----------------|-------------------|--------------------|----------------------|--------------|----------------|| SASRec | 0.341 +/- 0.008 | 0.412 +/- 0.011 | 0.318 +/- 0.006 | 0.397 +/- 0.009 | 0.276 +/- 0.007 | 0.358 +/- 0.010 || BERT4Rec | 0.353 +/- 0.007 | 0.421 +/- 0.010 | 0.327 +/- 0.005 | 0.402 +/- 0.008 | 0.284 +/- 0.006 | 0.364 +/- 0.009 || DDPG-Rec | 0.347 +/- 0.009 | 0.428 +/- 0.012 | 0.321 +/- 0.007 | 0.408 +/- 0.010 | 0.279 +/- 0.008 | 0.366 +/- 0.011 || SAC-Rec | 0.358 +/- 0.008 | 0.439 +/- 0.011 | 0.334 +/- 0.006 | 0.418 +/- 0.010 | 0.291 +/- 0.007 | 0.371 +/- 0.010 || SlateQ | 0.362 +/- 0.009 | 0.443 +/- 0.012 | 0.339 +/- 0.007 | 0.421 +/- 0.011 | 0.293 +/- 0.008 | 0.374 +/- 0.011 || FeedRec | 0.367 +/- 0.008 | 0.461 +/- 0.011 | 0.343 +/- 0.006 | 0.434 +/- 0.010 | 0.298 +/- 0.007 | 0.382 +/- 0.010 || RLUR | 0.371 +/- 0.008 | 0.473 +/- 0.010 | 0.348 +/- 0.006 | 0.442 +/- 0.009 | 0.301 +/- 0.007 | 0.391 +/- 0.010 || **ATD-DRL (ours)** | **0.388 +/- 0.007** | **0.508 +/- 0.009** | **0.364 +/- 0.005** | **0.476 +/- 0.008** | **0.317 +/- 0.006** | **0.418 +/- 0.009** | ATD-DRL achieves an average NDCG@10 improvement of 4.7% and a counterfactualretention improvement of 7.3% over the strongest baseline. The improvementreaches 11.4% in the late segment of long sessions, statistically significantat p < 0.01. ## Getting Started Unpack both archives into the same working directory: ```bashunzip atd_drl_code.zipunzip atd_drl_data.zip``` Install dependencies: ```bashcd atd_drl_codepip install -r requirements.txt``` Run all experiments and render every figure: ```bashpython main.py --mode evaluate --dataset all``` Train on a real dataset (link the data package into the expected locationfirst): ```bashmkdir -p data/rawln -s ../../atd_drl_data/kuairec data/raw/kuairecln -s ../../atd_drl_data/microvideo data/raw/microvideoln -s ../../atd_drl_data/mind data/raw/mind python main.py --mode train --config configs/config_kuairec.yaml``` ## Environment - Python 3.9 or above- PyTorch 2.0 or above- NumPy, Pandas, scikit-learn, SciPy- Matplotlib, Seaborn- tqdm, PyYAML ## Directory Reference ### Code package (`atd_drl_code/`) - `main.py` -- entry point for evaluation, training, and sanity checks- `configs/` -- YAML configuration files, one per dataset- `data/` -- data loader and download instructions- `models/` -- ATD-DRL model, drift-aware encoder, discount network, actor, twin-Q critic, and eight baselines- `training/` -- Actor-Critic-gamma cooperative trainer and loss functions- `evaluation/` -- ranking metrics, doubly robust estimator, session-segment evaluator- `experiments/` -- scripts for the four experiment groups (main, ablation, robustness, user profile)- `visualization/` -- figure rendering- `utils/` -- non-stationary MDP formalisation, seeding, config loading, Polyak averaging ### Data package (`atd_drl_data/`) - `kuairec/interactions.csv` -- KuaiRec interaction log- `microvideo/interactions.csv` -- MicroVideo-1.7M interaction log- `mind/interactions.csv` -- MIND-Short interaction log- `results/` -- eight CSV / JSON tables covering Tables 3-8 and the Figure 11 heat-map grid- `figures/` -- eleven rendered figures (300 DPI PNG) ## License Released under the MIT License. See the LICENSE file if present, or refer tothe standard MIT terms. ## Contact For questions about the implementation or reported results, please open anissue on the repository.



