遇见数据集

PRIME-CVD Data Asset 1: DAG-Simulated Cardiovascular Risk Cohort for Medical Informatics Education

收藏
Figshare2026-02-23 更新2026-04-28 收录
官方服务:

资源简介:

OverviewPRIME-CVD Data Asset 1 is a directed acyclic graph (DAG)-simulated cohort of 50,000 synthetic individuals designed to reproduce realistic demographic structure, socioeconomic gradients, cardiometabolic risk factor distributions, and clinically plausible five-year cardiovascular disease (CVD) incidence consistent with contemporary Australian primary prevention populations [1]. The dataset encodes established epidemiologic relationships among age, socioeconomic disadvantage (IRSD), behavioural risk factors, chronic disease states, biomarkers, and time-to-event outcomes within a transparent, parametrically specified causal framework.SecurityAll individuals are simulated entirely de novo using a fully parameterised DAG configured from publicly available epidemiologic summaries (e.g., Australian Institute of Health and Welfare, Australian Bureau of Statistics, and peer-reviewed literature). No patient-level electronic medical record (EMR) data were used in model construction, and no machine learning generative models (e.g., GANs, diffusion models, or large language models) were trained on real clinical data.Because the simulation is entirely mechanism-based rather than data-trained, there is no membership inference risk, no residual linkage risk, and no possibility of re-identification. None of the synthetic individuals correspond to real-world patients, and the dataset contains no direct identifiers, quasi-identifiers, or protected health information.Educational FocusDespite being fully simulated, the dataset preserves realistic subgroup imbalance and clinically meaningful risk gradients, enabling applied training in epidemiology, medical informatics, and health data science without governance barriers.PRIME-CVD Data Asset 1 is suitable for instruction in:Cox proportional hazards modelling (survival analysis in epidemiology)Risk prediction model development and calibration assessmentClassification metrics (precision, recall, F1 score) for statistical interpretationDimensionality reduction techniques (e.g., t-SNE) for data visualisationDemographic and socioeconomic stratification for health policy analysisFairness-aware modelling and subgroup performance evaluationThis environment allows learners to develop analytic workflows and methodological competence prior to working with governed clinical datasets.Reference[1] Kuo NI-H, et al. Estimating 5-year absolute risk of cardiovascular disease using routinely collected electronic medical records from Australian general practices. Heart. 2025.Synthetic Cohort Characteristics (N = 50,000)Age (years)Mean (SD): 49.71 (12.37)Median [IQR]: 49.63 [41.33, 58.09]Range: 18.0–90.0IRSD Quintile DistributionQ1: 21.28% (Most disadvantaged)Q2: 16.11%Q3: 23.88%Q4: 16.99%Q5: 21.74% (Least disadvantaged)Smoking StatusNon-smoker: 73.14%Ex-smoker: 16.72%Current smoker: 10.13%Chronic Disease PrevalenceDiabetes mellitus: 7.43%Chronic Kidney Disease (CKD): 0.680%Atrial Fibrillation (AF): 0.720%Body Mass Index (BMI, kg/m²)Mean (SD): 28.33 (5.03)Median [IQR]: 28.33 [24.92, 31.73]Range: 15.0–52.76Systolic Blood Pressure (SBP, mmHg)Mean (SD): 123.31 (16.10)Median [IQR]: 123.14 [112.39, 134.10]Range: 55.85–187.79Estimated Glomerular Filtration Rate (eGFR, mL/min/1.73m²)Mean (SD): 82.77 (6.09)Median [IQR]: 82.94 [79.22, 86.66]Range: 37.00–104.65Haemoglobin A1c (HbA1c, %)Mean (SD): 4.79 (0.93)Median [IQR]: 4.66 [4.24, 5.12]Range: 2.23–12.71Cardiovascular OutcomesOverall 5-year CVD event rate: 4.02%Mean follow-up time: 4.80 years

## 概述 PRIME-CVD数据资产1是一项基于有向无环图(directed acyclic graph, DAG)模拟的队列数据集,包含50000名合成个体,旨在复现符合当代澳大利亚一级预防人群特征的真实人口结构、社会经济梯度、心脏代谢危险因素分布,以及具备临床合理性的5年心血管疾病(cardiovascular disease, CVD)发病风险[1]。该数据集在透明的参数化因果框架内,编码了年龄、社会经济劣势(Index of Relative Socio-economic Disadvantage, IRSD)、行为危险因素、慢性疾病状态、生物标志物以及事件发生时间结局之间已确立的流行病学关联。 ## 数据安全 所有个体均通过完全参数化的DAG从头模拟生成,该DAG的配置基于公开可得的流行病学汇总数据(如澳大利亚健康与福利研究所、澳大利亚统计局以及同行评议文献)。模型构建过程未使用任何患者级电子病历(electronic medical record, EMR)数据,也未在真实临床数据上训练任何机器学习生成模型(如生成对抗网络(generative adversarial networks, GANs)、扩散模型或大语言模型(large language model, LLM))。 由于该模拟完全基于机制而非数据训练,因此不存在成员推断风险、残留关联风险,也无法实现个体重识别。所有合成个体均不对应真实世界患者,数据集也未包含直接标识符、准标识符或受保护健康信息。 ## 教学应用重点 尽管该数据集为完全模拟生成的数据,但仍保留了真实的亚组不均衡性与具有临床意义的风险梯度,可用于流行病学、医学信息学与健康数据科学领域的实操培训,且无需受治理临床数据所需的审批障碍。 PRIME-CVD数据资产1适用于以下教学场景: 1. Cox比例风险建模(流行病学领域的生存分析) 2. 风险预测模型开发与校准评估 3. 用于统计解读的分类指标(精确率、召回率、F1分数) 4. 用于数据可视化的降维技术(如t分布邻域嵌入(t-distributed stochastic neighbor embedding, t-SNE)) 5. 用于卫生政策分析的人口与社会经济分层分析 6. 公平感知建模与亚组性能评估 该环境可让学习者在接触受治理的临床数据集之前,就能够开发分析工作流并提升方法学应用能力。 ## 参考文献 [1] Kuo NI-H, 等. 基于澳大利亚全科诊所常规收集的电子病历估算5年心血管疾病绝对风险. 心脏(Heart). 2025. ## 合成队列特征(N=50000) ### 年龄(岁) 均值(标准差):49.71(12.37);中位数[四分位距]:49.63[41.33, 58.09];范围:18.0~90.0 ### IRSD五分位分布 Q1:21.28%(最劣势组);Q2:16.11%;Q3:23.88%;Q4:16.99%;Q5:21.74%(最优势组) ### 吸烟状态 非吸烟者:73.14%;既往吸烟者:16.72%;当前吸烟者:10.13% ### 慢性疾病患病率 糖尿病:7.43%;慢性肾脏病(chronic kidney disease, CKD):0.680%;心房颤动(atrial fibrillation, AF):0.720% ### 身体质量指数(body mass index, BMI,kg/m²) 均值(标准差):28.33(5.03);中位数[四分位距]:28.33[24.92, 31.73];范围:15.0~52.76 ### 收缩压(systolic blood pressure, SBP,mmHg) 均值(标准差):123.31(16.10);中位数[四分位距]:123.14[112.39, 134.10];范围:55.85~187.79 ### 估算肾小球滤过率(estimated glomerular filtration rate, eGFR,mL/min/1.73m²) 均值(标准差):82.77(6.09);中位数[四分位距]:82.94[79.22, 86.66];范围:37.00~104.65 ### 糖化血红蛋白(haemoglobin A1c, HbA1c,%) 均值(标准差):4.79(0.93);中位数[四分位距]:4.66[4.24, 5.12];范围:2.23~12.71 ### 心血管结局 总体5年CVD事件发生率:4.02%;平均随访时间:4.80年

创建时间:
2026-02-23
二维码
社区交流群
二维码
科研交流群
商业服务