遇见数据集

HUNTINGTON'S DISEASE ENRICHED COHORT

收藏
Zenodo2026-09-15 更新2026-10-01 收录
官方服务:

资源简介:

This Zenodo record contains a free, evaluation-grade sample of 10,000 synthetic patient records — a genuine, unmodified subset of the full production dataset, not a separate re-generation. It is provided so that researchers, ML engineers, and clinical data scientists can inspect the schema, phenotypic depth, and biomarker architecture before licensing the complete asset. The full commercial dataset — 5,000,000 synthetic patients, 11-year longitudinal depth (2016–2026), complete multi-ethnic population stratification, and the full pedigree network — is a licensed product and is not included in this record. Full dataset & enterprise licensing: https://sentineldata.com.ua/dataset/huntington-s-disease-cohort-dataset Abstract This dataset is a synthetic, privacy-safe electronic health record (EHR) cohort standardized to the OMOP Common Data Model (CDM) v5.4, built to address a persistent problem in rare-disease machine learning: the near-total absence of large, structurally complete, freely shareable datasets for orphan neurodegenerative conditions. The cohort is enriched around Huntington's Disease (HD), combining longitudinal clinical trajectories, genetic biomarker assertions (CAG repeat length), multi-generational pedigree graphs, and a realistic comorbidity background. No real patient data of any kind is contained in this file — every record is algorithmically generated. Clinical & Scientific Background Huntington's disease is an autosomal-dominant neurodegenerative disorder caused by an expanded CAG trinucleotide repeat in the HTT gene on chromosome 4p16.3, producing a toxic polyglutamine tract in the huntingtin protein and progressive striatal and cortical degeneration [1,7]. Clinically it presents as a triad of chorea, cognitive decline, and psychiatric disturbance, typically manifesting between 30 and 50 years of age, with juvenile-onset forms appearing in carriers of markedly longer repeats [6,7]. Repeat-length genetics follow defined clinical thresholds: alleles of 27–35 repeats are intermediate (unstable in transmission but not disease-causing in the carrier); 36–39 repeats are classified as reduced-penetrance; and 40 or more repeats are fully penetrant, virtually guaranteeing disease within a normal lifespan [1,2]. This threshold structure is the biological backbone against which any HD genetic dataset should be evaluated. Epidemiologically, HD is a genuinely rare disease. Pooled global prevalence is estimated at 3.9–4.9 per 100,000 persons, rising to 5.6–8.6 per 100,000 in populations of European descent (Europe, North America, Oceania) [3,4,5]. That scarcity is precisely why real-world HD datasets large enough for robust ML work are almost impossible to assemble without pooling many clinical sites — and why synthetic augmentation has scientific and commercial value. Why Synthetic Data — The Rare-Disease Gap Three structural problems make HD, and orphan diseases generally, a poor fit for conventional EHR data science: Prevalence is too low for statistical power. At roughly 5–8 cases per 100,000, a random hospital extract of even 1 million patients yields on the order of 50–80 HD cases — too few to train or validate most supervised models. Real HD registries are access-restricted. Data with genetic testing results and pedigree information is highly identifying; legitimate registries (e.g. Enroll-HD) require credentialed access and cannot be freely redistributed. Prodromal and family-structure data are the hardest signal to obtain. The features researchers most want — psychiatric prodrome, inheritance vectors, biomarker trajectories — are exactly the features most restricted in real data due to re-identification risk. Synthetic generation removes the access barrier entirely: no consent, no IRB, no re-identification risk, while preserving the statistical shape of the disease. Cohort Construction & Sampling Methodology This 10,000-patient sample uses stratified enrichment sampling, not a simple random draw from the full 5,000,000-patient population. A random draw at true prevalence (~8 per 100,000) would return close to zero HD-positive patients in a sample this size — useless for evaluation purposes. Instead, this sample is deliberately weighted to guarantee representation of: confirmed HD patients with full symptom trajectories, their pedigree-linked relatives (parent/child edges), the genetic biomarker panel, a realistic non-HD background population for contrast and noise modeling. The full 5,000,000-patient licensed dataset, by contrast, keeps HD at its真 real-world epidemiological rate (~0.008% / 8 per 100,000) embedded in a full-scale, non-enriched background population — matching real hospital-network case density rather than an artificially inflated one. Data Model — OMOP CDM v5.4 Standardized to OMOP CDM v5.4, maintained by OHDSI (Observational Health Data Sciences and Informatics), the current supported CDM release and the de facto standard for federated observational health research [7]. Tables included in this sample: Table Purpose PERSON Demographics: birth date, gender, race, ethnicity OBSERVATION_PERIOD Temporal coverage window per patient VISIT_OCCURRENCE Inpatient / outpatient / ER encounters CONDITION_OCCURRENCE SNOMED-CT diagnoses, staged by disease progression DRUG_EXPOSURE RxNorm-coded prescriptions MEASUREMENT Lab values and the CAG repeat genetic assay (LOINC/OMOP 4152011) FACT_RELATIONSHIP Parent-of / child-of pedigree edges Standardized Concept Dictionary (selected) Domain Concept Concept ID Vocabulary Condition Huntington's Disease 434221 SNOMED-CT Condition Chorea 376304 SNOMED-CT Condition Dysphagia 4184643 SNOMED-CT Condition Cachexia 433736 SNOMED-CT Condition Depressive Disorder 440383 SNOMED-CT Condition Injury / Trauma 432791 SNOMED-CT Condition Essential Hypertension 320128 SNOMED-CT Condition Type 2 Diabetes Mellitus 201820 SNOMED-CT Drug Tetrabenazine 10454 RxNorm Measurement CAG Repeat Length 4152011 LOINC/OMOP Relationship Parent of 40485452 OMOP Relationship Child of 41436030 OMOP Genetic Biomarker Sub-Model — CAG Repeat Assay The MEASUREMENT domain includes a CAG repeat length assay against real clinical thresholds: under 36 repeats is non-pathogenic/intermediate, 36–39 is reduced-penetrance, and 40 or more is fully penetrant [1,2]. Values in the sample span the intermediate/reduced-penetrance zone through highly-expanded, fully-penetrant alleles — consistent with testing extending across the pedigree (at-risk relatives as well as confirmed probands), not only diagnosed cases. Family Pedigree Graph FACT_RELATIONSHIP records encode a bidirectional kinship graph (parent-of / child-of edge pairs), enabling multi-generational inheritance-vector analysis and graph neural network training on hereditary transmission patterns — a structure almost never available in de-identified real-world EHR extracts. File Format Delivered as Apache Parquet (snappy compression), directly queryable with DuckDB, Polars, or PySpark without loading the full file into memory. Column-level types and schema definitions are documented in the accompanying technical file. Licensing This sample is distributed under Creative Commons Attribution 4.0 International (CC BY 4.0) — free for any use, commercial or non-commercial, with attribution. The full 5,000,000-patient dataset is a separate commercial license obtained via sentineldata.com.ua. Intended Use & Limitations This is synthetic data generated for machine learning development, algorithm validation, and educational use. It does not describe real patients and must not be used to make or support any individual clinical decision. Statistical properties are calibrated against published epidemiological and genetic literature but have not been independently peer-reviewed; users conducting downstream research should validate assumptions against primary sources cited below. Citation Sentinel Data. Huntington's Disease (HD) Enriched Synthetic Cohort — Free Sample (N=10,000), OMOP CDM v5.4. 2026. Full dataset: https://sentineldata.com.ua/dataset/huntington-s-disease-cohort-dataset References [1] American College of Medical Genetics and Genomics. Standards and Guidelines for Clinical Genetics Laboratories: Huntington Disease. Genetics in Medicine, 2014. https://www.nature.com/articles/gim2014146[2] GeneReviews (NCBI Bookshelf). Huntington Disease. NBK1305. https://www.ncbi.nlm.nih.gov/books/NBK1305/[3] Medina A, et al. Prevalence and Incidence of Huntington's Disease: An Updated Systematic Review and Meta-Analysis. Movement Disorders, 2022. PMID: 36161673.[4] Huntington Study Group. How Many People Have Huntington Disease? 2024. https://huntingtonstudygroup.org/hd-insights/how-many-people-have-huntington-disease/[5] Rare Disease Advisor. Huntington Disease Epidemiology. 2025. https://www.rarediseaseadvisor.com/disease-info-pages/huntington-disease-epidemiology/[6] Rare Disease Advisor. Addressing the Underlying Causes of Huntington Disease. 2026. https://www.rarediseaseadvisor.com/insights/addressing-underlying-causes-huntington-disease/[7] OHDSI. OMOP Common Data Model v5.4. https://github.com/OHDSI/CommonDataModel

提供机构:
Zenodo
创建时间:
2026-09-15
二维码
社区交流群
二维码
科研交流群
商业服务