De-identified dataset and analysis code for "Development of a Fetal Congenital Heart Disease Risk Prediction Model Based on Machine Learning Algorithms"
收藏资源简介:
This record contains the de-identified minimal analysis dataset and reproduction code for the study "Development of a Fetal Congenital Heart Disease Risk Prediction Model Based on Machine Learning Algorithms". The dataset includes 25,022 singleton pregnancies (training set: 15,748 cases, including 430 fetal congenital heart disease [CHD] cases and 15,318 non-CHD controls; external validation set: 9,274 cases from a different NIPT-Plus sequencing platform, including 279 CHD cases and 8,995 controls). For each pregnancy, the 19 model input features (NIPT chromosome Z-score-derived indicators, fetal fraction, first- and second-trimester serum biomarkers, nuchal translucency, biparietal diameter, and maternal age), the CHD outcome, and CHD severity are provided. All direct identifiers were removed prior to deposition (names, hospital numbers, sample barcodes, ID numbers, contact details, addresses, birth dates, and all sampling/reporting/delivery dates), and record IDs were randomly re-assigned. The chromosome Z-values included are aggregated, sample-level summary statistics routinely reported by clinical non-invasive prenatal testing laboratories; no raw sequencing data, sequence reads, or genotype-level data are included. File access is restricted; access can be requested via the "Request access" button. See README.md for the full data dictionary and reproduction instructions.



