Synthetic Longitudinal Mental Health Dataset (1M Sample; PHQ-9, GAD-7, PCL-5)
收藏资源简介:
This record releases a research-grade synthetic longitudinal mental health dataset designed for methodological research, benchmarking, and education in contexts where access to real mental-health microdata is ethically or legally restricted. The dataset simulates weekly longitudinal observations for individuals, capturing symptom dynamics related to depression, anxiety, and post-traumatic stress disorder (PTSD) using established psychometric instruments with item-level ordinal response generation: PHQ-9 (9 items and total score) GAD-7 (7 items and total score) PCL-5 (20 items and total score), including PTSD symptom cluster subscales B, C, D, and E The publicly released archive (mh_10m_sample_1m.zip) contains a 1,000,000-row Parquet sample suitable for open distribution and rapid experimentation. This sample is derived from a larger synthetic simulation comprising 10,617,183 total longitudinal observations, as reported in the accompanying validation summary. Dataset contents Each observation includes multiple classes of variables commonly encountered in applied mental-health research: Psychometric variables Item-level responses and total scores for PHQ-9, GAD-7, and PCL-5 PTSD cluster-level subscale scores (B/C/D/E) Demographic and contextual covariates Age, gender, ethnicity (coded categories) Socioeconomic status (SES) indicators Urban/rural indicator Education level, profession, and marital status Clinical history and risk proxies Family mental-health history indicator Chronic illness indicators (including diabetes, cardiovascular disease, gastrointestinal conditions, and thyroid/hormonal conditions) Trauma exposure proxy Behavioral and contextual measures Sleep, stress, social support, and physical activity measures Wearable-style proxies Step counts Screen time Heart-rate variability (HRV) Study design and reporting realism Recruitment source indicators and sampling weights Longitudinal attrition indicators (dropout and withdrawal) Simulated response distortion indicators (e.g., underreporting flags) Measurement and lab-bias indicators Treatment and longitudinal dynamics Time-varying treatment indicator Treatment adherence intensity Repeated observations per individual enabling longitudinal prediction and causal analysis Synthetic generation procedures The data generation process was explicitly designed to resemble real observational mental-health datasets by modeling the mechanisms that introduce statistical complexity in applied research: Latent symptom trajectories over timeDepression, anxiety, and PTSD latent states evolve longitudinally with temporal dependence and stochastic perturbations, producing realistic within-individual variability across repeated observations. Item-level psychometric modelingQuestionnaire items are generated from latent states using graded ordinal response logic rather than sampling aggregate scores directly, preserving scale structure and high internal consistency. Realistic missingness and response behaviorItem-level nonresponse is simulated at empirically plausible rates, with additional response-behavior mechanisms (e.g., underreporting and response distortion) included to reflect measurement error commonly observed in self-report mental-health data. Selection, recruitment, and attrition mechanismsParticipation is shaped by recruitment and selection processes, followed by longitudinal attrition, producing incomplete and unbalanced panels typical of real cohort studies. Time-varying treatment and confoundingTreatment assignment varies over time and is confounded by evolving symptom severity and covariates, enabling evaluation of causal inference methods under realistic non-random treatment assignment. Validation and benchmarking Validation and benchmarking were conducted on a random 80,000-row evaluation sample, with results provided in the accompanying files (metrics.csv, DIF outputs): Total simulated dataset size: 10,617,183 rows Item-level missingness rates: PHQ-9 ≈ 8.19% GAD-7 ≈ 8.11% PCL-5 ≈ 11.10% Internal consistency (Cronbach’s α): PHQ-9 ≈ 0.965 GAD-7 ≈ 0.954 PCL-5 ≈ 0.980 Predictive benchmarking: Severe depression classification performance: AUC ≈ 0.93 Imputation benchmarking: Masked PHQ-1 recovery: RMSE ≈ 0.48 Causal benchmarking: Doubly robust average treatment effect (ATE) estimation Comparison with a simulated counterfactual subset (reported in metrics.csv) Measurement bias and invariance screening outputs are included: Mantel–Haenszel DIF results (DIF_MH_PHQ_gender.csv) Logistic regression DIF results (DIF_LR_PHQ_gender.csv) Files included in this record mh_10m_sample_1m.zip — 1,000,000-row Parquet dataset metrics.csv — validation and benchmark summary DIF_MH_PHQ_gender.csv — Mantel–Haenszel DIF results DIF_LR_PHQ_gender.csv — logistic regression DIF results eval_log.txt — evaluation log (transparency artifact) Scientific scope and intended use This dataset is explicitly intended for research use, including: Statistical and machine-learning benchmarking Missing-data and imputation research Fairness and measurement-bias (DIF) analysis Longitudinal modeling Causal inference and treatment-effect estimation Teaching and demonstration involving sensitive mental-health data structures Limitations and appropriate use This dataset is fully synthetic and does not represent real patient populations.It is intended for methods research and analytical benchmarking, not for: clinical decision-making or diagnosis, estimation of real-world prevalence, incidence, or risk, drawing substantive clinical or epidemiological conclusions about real populations. All statistical relationships are simulated under explicit modeling assumptions and are designed to support evaluation of analytical methods rather than inference about real-world mental health outcomes.



