遇见数据集

Fabry Disease Case Report Dataset: Structured clinical data from 40 open-access publications

收藏
Zenodo2026-02-27 更新2026-05-29 收录
官方服务:

资源简介:

Fabry Disease Case Report Dataset Structured clinical data from open-access case reports, extracted via AI pipeline Authors Adrian Michalski (ORCID: 0000-0003-0800-5394) Abstract This dataset contains structured clinical phenotype data for 61 patients with Fabry Disease (Anderson-Fabry Disease, OMIM #301500), extracted from 40 open-access case reports indexed in PubMed Central. Data was extracted using the Silene pipeline (Silene Systems), which combines large language model-based information extraction with automated clinical validation and normalization against standard ontologies (HPO, LOINC). The dataset captures demographics, genetic variants, symptoms, laboratory values, phenotype classifications, and disease staging for each patient. All source publications are open-access and Creative Commons licensed. Methodology Data was extracted using a multi-stage AI pipeline (Silene, by Silene Systems): Search - PubMed Central queried for Fabry Disease case reports using E-utilities API Screen - AI-based relevance screening to filter case reports with extractable patient-level data Extract - Structured extraction of patient demographics, genetics, symptoms, and lab values from full-text articles using Claude (Anthropic) with domain-specific prompts Validate - Automated clinical validation checking internal consistency (e.g., genotype-phenotype concordance, age plausibility, lab value ranges) Normalize - Symptom names mapped to HPO terms, lab values standardized to canonical names and units, phenotype classification harmonized Each record includes an extraction_confidence score (model self-assessment) and a data_completeness score (proportion of fields populated). Data Dictionary patients.csv / patients.json Field Type Description case_id string Unique patient identifier (format: FAB-NNN) sex string Patient sex: "male" or "female" age_at_diagnosis integer Age in years when Fabry Disease was diagnosed age_at_symptom_onset integer Age in years when first symptoms appeared genetic_variant string GLA gene variant in HGVS coding DNA notation (e.g., c.679C>T) protein_change string Predicted protein change in HGVS notation (e.g., p.R227X) zygosity string Hemizygous (males), heterozygous (females), or homozygous phenotype string Clinical phenotype: "classic", "late-onset", "cardiac", "renal", or "unknown" disease_stage string Disease progression: "early", "moderate", "advanced", or "unknown" extraction_confidence float AI extraction confidence score (0.0-1.0) data_completeness float Proportion of fields successfully extracted (0.0-1.0) diagnosis_summary string Free-text clinical summary of the case JSON-only nested fields: Field Type Description symptoms array List of symptom objects (see symptoms.csv) lab_values array List of lab value objects (see lab_values.csv) source object Source publication metadata (pmcid, pmid, doi, title, journal, publication_date, license) symptoms.csv Long-format table of clinical symptoms per patient. Field Type Description case_id string Patient identifier (FK to patients) symptom_name string Normalized symptom name (mapped to HPO where possible) present boolean Whether the symptom is present (True) or explicitly absent (False) details string Additional clinical details, timing, severity lab_values.csv Long-format table of laboratory and clinical measurements. Field Type Description case_id string Patient identifier (FK to patients) lab_name string Normalized lab test name value string Measured value (numeric or descriptive) unit string Unit of measurement timepoint string When the measurement was taken (e.g., "at diagnosis", "6 months post-ERT") context string Additional context (e.g., "reference value < 1.11 ng/mL") publications.csv Source publication metadata with enriched identifiers. Field Type Description pmcid string PubMed Central identifier (primary key) pmid string PubMed identifier (enriched via NCBI ID Converter) doi string Digital Object Identifier (enriched via NCBI ID Converter) title string Publication title journal string Journal name publication_date string Publication date (YYYY-MM-DD) license string Creative Commons license type (e.g., CC BY, CC BY-NC) Dataset Statistics Metric Value Total patients 61 Source publications 40 Male / Female 36 / 22 Mean age at diagnosis 47.6 years Mean age at symptom onset 31.2 years Mean extraction confidence 0.90 Mean data completeness 0.61 Field Coverage Field Coverage sex high age_at_diagnosis moderate-high genetic_variant moderate-high protein_change moderate phenotype high (classified for all) symptoms high (variable count per patient) lab_values moderate (variable count per patient) age_at_symptom_onset moderate zygosity moderate Known Limitations Case report bias - Case reports are published for atypical, severe, or diagnostically challenging presentations. This dataset is not representative of the general Fabry Disease population. AI extraction errors - Despite validation, AI-extracted data may contain errors. The extraction_confidence score provides a per-record quality indicator but is not a guarantee of accuracy. Sparse fields - Not all fields are reported in every source publication. data_completeness reflects the proportion of fields populated, but missing data should be treated as unreported, not absent. Temporal heterogeneity - Source publications span multiple years and reflect different clinical practices, diagnostic criteria, and treatment availability. Normalization artifacts - Symptom and lab value names have been normalized to canonical forms. Some clinical nuance may be lost in this process; refer to the details and context fields for original descriptions. Single-gene focus - All patients have Fabry Disease caused by GLA gene variants. The dataset structure could support other lysosomal storage disorders but currently only contains Fabry Disease data. License The structured data in this dataset is released under CC BY 4.0 (Creative Commons Attribution 4.0 International). All source publications are open-access and published under Creative Commons licenses (CC BY, CC BY-NC, CC BY-NC-ND, CC BY-NC-SA). The extracted factual data (clinical observations, measurements, genetic variants) is not subject to copyright per the fact/expression dichotomy. The structured compilation is an original work licensed CC BY 4.0. Suggested Citation Michalski, A. (2026). Fabry Disease Case Report Dataset: Structured clinical data from 40 open-access publications. Zenodo. https://doi.org/10.5281/zenodo.18662232 Extraction Pipeline This dataset was generated using Silene Systems (https://github.com/amichalski/silene-case-study), an AI-powered pipeline for extracting structured clinical data from medical case reports. File Manifest File Format Description patients.json JSON Full dataset with nested symptoms, lab values, and source metadata patients.csv CSV Flat patient demographics and genetics table symptoms.csv CSV Long-format symptom presence/absence per patient lab_values.csv CSV Long-format laboratory measurements per patient publications.csv CSV Source publication metadata with enriched identifiers README.md Markdown This data descriptor Contact For questions about this dataset, please contact the author via ORCID.

提供机构:
Zenodo
创建时间:
2026-02-16
二维码
社区交流群
二维码
科研交流群
商业服务