Synthetic dataset and analysis code for AI-enabled digital phenotyping in Internet Gaming Disorder risk stratification
收藏资源简介:
This record contains the fully synthetic reference dataset, revised Python analysis code, supplementary/result tables, and analysis outputs for the simulation study “AI-Enabled Digital Phenotyping for Personalized Risk Stratification in Internet Gaming Disorder: A Privacy-Preserving Simulation Study.” The reference dataset includes 1,000 virtual user profiles generated using random seed 42. Four aggregated telemetry features were simulated: average session duration, sessions per week, Late-Night Index, and application-switching rate. The simulated elevated-risk outcome was generated from a known weighted latent risk function, with the top 20% assigned to the clean elevated-risk class and balanced stochastic label-noise injection applied to approximate imperfect ground truth. The revised analysis includes evaluation of Random Forest, Logistic Regression, and Gradient Boosting models using a stratified 80:20 train–test split; stratified five-fold cross-validation within the training set; Random Forest feature-importance and calibration analyses; feature-correlation analysis; sensitivity analyses across label-noise levels from 0% to 20%; and an across-seed robustness analysis based on 200 independent synthetic realizations under the primary 5% label-noise condition. Seed 42 is retained as the reference reproducible realization reported in the primary manuscript analyses. The simulation is intended as a controlled methodological test environment. The all-feature versus playtime-only comparison represents an internal-consistency assessment of the pre-specified synthetic signal structure and should not be interpreted as empirical evidence that the simulated telemetry features are superior predictors of real-world Internet Gaming Disorder. No real participants, patient records, identifiable behavioral logs, raw keystrokes, screenshots, message content, or other content-level behavioral data were used. All records are fully synthetic and are provided to support transparency and reproducibility of the manuscript and its revised analyses.



