遇见数据集

Emory White Matter Hyperintensity Dataset

收藏
Zenodo2026-06-05 更新2026-05-26 收录
官方服务:

资源简介:

The Emory White Matter Hyperintensity (WMH) Dataset is a retrospectively curated collection of defaced clinical brain MRI scans with expert manual WMH annotations, developed to support robust evaluation of WMH segmentation methods under real-world clinical heterogeneity. The dataset comprises 195 routine clinical MRI examinations acquired between 2006 and 2022 from 71 distinct scanners, spanning multiple vendors, magnetic field strengths, and acquisition protocols. This diversity reflects real-world clinical variability and enables systematic assessment of model generalizability and robustness. White matter hyperintensities were manually delineated by an experienced rater and reviewed under neuroradiologist supervision. The dataset includes fixed training and test splits, together with standardized participant- and scanner-level metadata, to facilitate reproducible benchmarking and fair comparison of WMH segmentation algorithms. All MRI images have been defaced, and no direct personal identifiers are included. Data Access Data access is managed through the Alzheimer's Disease Data Initiative (ADDI) AD Workbench. All requests for access to the Emory WMH Dataset must be submitted via: https://discover.alzheimersdata.org/catalogue/datasets/9f93737c-ecf2-4e62-a748-d626f6e84ec1 Related publication This dataset accompanies the peer-reviewed article: Benchmark white matter hyperintensity segmentation methods fail on heterogeneous clinical MRI: A new dataset and deep learning-based solutions Wu J, Brown JD, Hu R, Edwards PJ, Levey AI, Lah JJ, Qiu D. Journal of Imaging Informatics in Medicine. https://doi.org/10.1007/s10278-025-01808-9 Processing Pipeline A pre-configured Docker image for the Emory Robust WMH Segmentation pipeline is available at:https://hub.docker.com/r/emorycn2l/emory_robust_wmh It provides a standardized, fully reproducible environment for robust, automated WMH segmentation on heterogeneous clinical MRI (FLAIR and T1-weighted images). The included nnU-Net model was trained on the Emory White Matter Hyperintensity Dataset. Dataset organization The dataset is organized into fixed training and test splits with one directory per participant: emory_wmh_dataset/├── README.md├── train/│ └── sub-XXXXXXXX/│ ├── original/│ ├── preprocessed/│ └── wmh.nii.gz├── test/│ └── sub-XXXXXXXX/│ ├── original/│ ├── preprocessed/│ └── wmh.nii.gz└── participants.tsv Each participant is assigned a pseudonymous identifier (sub-XXXXXXXX). Imaging data and annotations For each participant, the dataset includes: Preprocessed images Defaced, N4 bias-corrected FLAIR image Defaced, N4 bias-corrected T1-weighted image co-registered to FLAIR space Original images Defaced FLAIR and T1-weighted images in native acquisition space Defacing mask for T1-weighted image Affine transformation matrix mapping native T1-weighted space to FLAIR space (FSL FLIRT format) WMH annotation Binary WMH segmentation mask aligned to preprocessed FLAIR space (0 = background, 1 = WMH) Metadata Participant- and scanner-level metadata are provided in participants.tsv and may include: Demographics and clinical variables (e.g., age, sex, diagnosis, etiology, cognitive scores) CSF biomarkers (when available) Scanner information, including anonymized scanner identifiers and magnetic field strength Dataset split designation (train or test) These metadata support stratified analyses across demographic, clinical, and acquisition-related factors. Privacy and data protection All images are defaced to remove facial features. Scanner identifiers are pseudonymized using a salted one-way cryptographic transform, preventing re-identification while preserving scanner-level analytical utility.

提供机构:
Zenodo
创建时间:
2026-01-22
二维码
社区交流群
二维码
科研交流群
商业服务