遇见数据集

Dataset for Interpretable Machine Learning for Population-Level Tooth Loss Prediction

收藏
Zenodo2026-05-15 更新2026-05-26 收录
官方服务:

资源简介:

Reproducibility Package for the revised manuscript: Interpretable Machine Learning for Population-Level Tooth Loss Prediction Submitted to the Journal of Dental Research (JDR-26-0410, Under Review). Revision Context This package accompanies the revised manuscript submitted in response to peer review. The revision addressed all reviewer comments, including six substantive analytical additions that required new code, new outputs, and retraining of all models: Full-data retraining (Reviewer 1, Comment 3): The original 80/20 derivation split was replaced by training on 100% of BRFSS 2022 (N=433,772). Both the submitted 80/20 model and the revised full-data model are included for transparent comparison. Optimism-corrected performance (Reviewer 1, Comment 5): A 100-replicate bootstrap optimism correction was added for the full-data EBM on BRFSS 2022, reporting apparent, optimism, and optimism-corrected AUC and Brier score. Sample-size justification (Reviewer 1, Comment 7): Formal development (pmsampsize) and validation (pmvalsampsize) sample-size calculations were added following Riley et al. False-positive burden (Reviewer 1, Comment 9): A threshold operating-characteristics table was added reporting sensitivity, specificity, PPV, NPV, false-positive rate, and number needed to screen at clinically relevant probability thresholds (5%–50%) on both BRFSS 2024 and NHANES. BRFSS 2022 development performance (Reviewer 2, Comment 5): Apparent performance on the derivation dataset is now reported to enable comparison with studies that did not use external validation. Code and model sharing (Reviewers 1 and 2, Comments 8 and 4): This Zenodo package provides the complete reproducibility archive with code, trained models, and derived analytic data, fulfilling the open-data requirement. Additional manuscript-level revisions (retrospective temporal validation terminology, removal of "clinically deployable" claims, intended-use/monitoring/accountability section, abbreviation audit, supplement cross-references, non-inferiority margin justification, and descriptive Results language) are reflected in the revised manuscript text but do not require separate reproducibility artifacts. What This Package Contains This archive provides all analytic inputs, code, trained model artifacts, and final outputs needed to independently verify and reproduce the results reported in the revised manuscript. Data (8 files): Harmonized analytic CSV files derived from public-use CDC BRFSS 2022, BRFSS 2024, and NHANES 2015–2018 datasets, plus MICE-imputed Feather files from both the submitted pipeline (80/20 split) and the revised full-data pipeline. Models (6 artifacts): Both the submitted 80/20 EBM and the revised full-data EBM, locked Optuna hyperparameters, full-data MICE imputer, MICE feature list, and NHANES isotonic calibrator. Results: Optimism-corrected bootstrap estimates (100 replicates), sample-size justification (pmsampsize/pmvalsampsize with R script and pre-generated output), threshold operating characteristics at 5 probability cut-points × 2 validation datasets, old-versus-full-data EBM comparison with decision-gate summary, 7-model benchmark comparison (LR, RF, XGBoost, CatBoost, LightGBM, MLP, Stacked Ensemble), NHANES recalibration metrics (pre- and post-isotonic), and a consistency ledger mapping every number in the manuscript to its source CSV cell. Figures: Publication-ready Figure 1 (EBM feature importance), Figure 2 (ROC curves, calibration plots, decision curve analysis across 3 datasets), and Figure 3 (surveillance-to-clinical translation governance framework), with underlying plot data CSVs. Code: Two Python pipeline scripts covering MICE imputation, EBM training, benchmark evaluation, isotonic recalibration, bootstrap, threshold analysis, sample-size calculation, and figure generation. A verify_package.py script performs automated integrity checks (SHA-256 hashes, private-path scanning, and semantic numerical validation of key results). Documentation: Data dictionary with BRFSS/NHANES source-variable mappings and value coding for all 24 columns and 19 missingness indicators, data-source URLs, reproducibility notes, R session information, and a SHA-256 manifest for all 55 files. Reproduction Two scripted reproduction paths are provided (PowerShell for Windows, Bash for macOS/Linux): Quick verification from included model artifacts (~5 min): regenerates tables, figures, and the consistency ledger, then runs the automated verifier. Full rebuild from clean analytic CSVs (~2–4 hours): refits the MICE imputer on 100% of BRFSS 2022, retrains the EBM and all 7 benchmarks, reruns isotonic recalibration, recomputes 100-replicate bootstrap, and regenerates all outputs from scratch. Key Results Verified by the Package BRFSS 2022 full-data EBM apparent AUC: 0.8616 | Optimism-corrected AUC: 0.8604 BRFSS 2024 retrospective temporal validation AUC: 0.8627 | Brier: 0.0845 NHANES direct-transfer AUC: 0.7538 | Post-isotonic recalibration Brier: 0.1363 Development sample: N=433,772 (71,023 events, EPV=1,868) Software Requirements Python ≥ 3.10 (exact pinned versions in requirements_freeze.txt; minimum versions in requirements.txt) R with pmsampsize and pmvalsampsize packages (only required if rerunning the sample-size R script; pre-generated output is included) License MIT License. The included analytic files are derived from publicly available, de-identified CDC survey data (BRFSS and NHANES). Associated Publication Lam QT et al. Interpretable Machine Learning for Population-Level Tooth Loss Prediction. Journal of Dental Research. 2026. [Under Review]

修订版手稿可复现性数据包: 面向人群水平牙齿缺失预测的可解释机器学习 提交至《牙科研究杂志》(Journal of Dental Research,JDR-26-0410,处于审稿中)。 修订背景 本数据包随附针对同行评议修订后的手稿。本次修订已回应全部审稿意见,其中包含六项实质性分析补充工作,需新增代码、生成新输出并重新训练所有模型: 1. 全数据重训练(审稿人1,意见3):将原有的80/20推导集拆分替换为使用100%的行为风险因素监测系统(Behavioral Risk Factor Surveillance System,BRFSS)2022数据(样本量N=433,772)进行训练。为便于透明化对比,本数据包同时包含提交版的80/20模型与修订后的全数据模型。 2. 乐观校正性能(审稿人1,意见5):针对BRFSS 2022上的可解释提升机(Explainable Boosting Machine,EBM)新增100次重复自助法乐观校正,报告表观性能、乐观偏差与乐观校正后的受试者工作特征曲线下面积(AUC)与布里尔分数(Brier Score)。 3. 样本量合理性论证(审稿人1,意见7):参考Riley等的研究,新增正式的开发集(pmsampsize)与验证集(pmvalsampsize)样本量计算方法。 4. 假阳性负担分析(审稿人1,意见9):新增阈值操作特征表,报告在临床相关概率阈值(5%~50%)下,BRFSS 2024与美国国家健康与营养检查调查(National Health and Nutrition Examination Survey,NHANES)两个数据集上的灵敏度、特异度、阳性预测值、阴性预测值、假阳性率以及所需筛查人数。 5. BRFSS 2022开发集性能(审稿人2,意见5):新增推导数据集上的表观性能报告,以便与未使用外部验证的研究进行对比。 6. 代码与模型共享(审稿人1、2,意见8、4):本Zenodo数据包提供完整可复现性归档,包含代码、训练好的模型与衍生分析数据,满足开放数据要求。 手稿层面的其他修订(包括回顾性时间验证术语调整、移除“可临床部署”表述、新增用途/监测/问责章节、缩写校验、补充文献交叉引用、非劣效性界值合理性论证以及结果描述语言优化)已体现在修订后的手稿正文中,无需额外可复现性配套材料。 本数据包包含内容 本归档包含所有分析输入、代码、训练好的模型工件与最终输出,可独立验证并复现修订版手稿中报告的所有结果。 数据(共8个文件):源自美国疾病控制与预防中心(Centers for Disease Control and Prevention,CDC)公开使用的BRFSS 2022、BRFSS 2024与NHANES 2015–2018数据集的标准化分析CSV文件,以及来自提交版流水线(80/20拆分)与修订版全数据流水线的多重插补链式方程(Multiple Imputation by Chained Equations,MICE)填充后的Feather格式文件。 模型(共6个工件):包含提交版的80/20拆分EBM与修订后的全数据EBM、锁定的Optuna超参数、全数据MICE填充器、MICE特征列表以及NHANES等压校准器。 结果:包含100次重复自助法乐观校正估计、样本量合理性论证(含R脚本与预生成输出的pmsampsize/pmvalsampsize分析)、2个验证数据集上5个概率截断点的阈值操作特征、新旧全数据EBM对比(含决策门汇总)、7模型基准对比(包括逻辑回归、随机森林、极端梯度提升、CatBoost、LightGBM、多层感知机与堆叠集成)、NHANES等压校准前后的重新校准指标,以及将手稿中所有数值映射至其来源CSV单元格的一致性台账。 图表:包含可直接用于发表的图1(EBM特征重要性)、图2(跨3个数据集的受试者工作特征曲线、校准图与决策曲线分析)以及图3(监测-临床转化治理框架),并附带对应绘图数据CSV文件。 代码:包含2个Python流水线脚本,涵盖MICE填充、EBM训练、基准模型评估、等压重新校准、自助法分析、阈值分析、样本量计算与图表生成。另有verify_package.py脚本可执行自动化完整性检查(包括SHA-256哈希校验、私有路径扫描与关键结果的语义数值验证)。 文档:包含BRFSS/NHANES源变量映射与所有24个列及19个缺失指示符的数值编码的数据字典、数据源URL、可复现性说明、R会话信息,以及所有55个文件的SHA-256清单。 复现流程 本数据包提供两种脚本化复现路径(Windows系统使用PowerShell,macOS/Linux系统使用Bash): - 快速验证(基于已包含的模型工件,耗时约5分钟):重新生成表格、图表与一致性台账,随后运行自动化验证脚本。 - 完整重建(基于干净的分析CSV文件,耗时约2~4小时):在100%的BRFSS 2022数据上重新拟合MICE填充器、重新训练EBM与全部7个基准模型、重新运行等压重新校准、重新计算100次重复自助法分析,并从头生成所有输出结果。 本数据包可验证的关键结果 BRFSS 2022全数据EBM表观AUC:0.8616 | 乐观校正后AUC:0.8604 BRFSS 2024回顾性时间验证AUC:0.8627 | 布里尔分数:0.0845 NHANES直接迁移AUC:0.7538 | 等压重新校准后布里尔分数:0.1363 开发集样本:N=433,772(事件数71,023,每变量事件数(Events Per Variable,EPV)=1,868) 软件要求 Python ≥ 3.10(精确锁定版本见requirements_freeze.txt;最低版本要求见requirements.txt) 需安装pmsampsize与pmvalsampsize包的R环境(仅当重新运行样本量计算R脚本时需要,预生成的输出结果已包含) 许可证 采用MIT许可证。本数据包包含的分析文件源自公开可获取的去标识化CDC调查数据(BRFSS与NHANES)。 关联出版物 Lam QT等. 面向人群水平牙齿缺失预测的可解释机器学习. 《牙科研究杂志》. 2026. [处于审稿中]

提供机构:
Zenodo
创建时间:
2026-05-15
二维码
社区交流群
二维码
科研交流群
商业服务