遇见数据集

FOMO60K

收藏
魔搭社区2026-07-13 更新2026-07-15 收录
官方服务:

资源简介:

# FOMO50K: Brain MRI Dataset for Large-Scale Self-Supervised Learning with Clinical Data ![fomo60k](https://cdn-uploads.huggingface.co/production/uploads/67fe896341c831de7c499f73/JW5iLH3NvpHi1SZJkwHMj.png) _Dataset paper preprint:_ **A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning**.<br/> [https://arxiv.org/pdf/2506.14432](https://arxiv.org/pdf/2506.14432). This repo contains FOMO50K containing 50K brain MRI scans. This dataset was released as FOMO60K and formed the basis for the **FOMO25: Foundation Model Challenge for Brain MRI** hosted at **MICCAI 2025**. See [fomo25.github.io](https://fomo25.github.io) and [this paper](https://arxiv.org/abs/2604.11679) for more information about the challenge. **New:** We also provide [FOMO300K](https://huggingface.co/datasets/FOMO-MRI/FOMO300K), a larger version with over 300K scans without co-registration and skull-stripping. ## Updates ``` April 3 2026: FOMO60K is now FOMO50K. MGH-Wild has changed license and is no longer part of the dataset. Instructions how how to reproduce FOMO60K will be updated soon. ``` ## Description FOMO50K is a large-scale dataset of brain MRI scans, including both clinical and research-grade scans. The dataset includes a wide range of sequences, including _T1, MPRAGE, T2, T2*, FLAIR, SWI, T1c, PD, DWI, ADC, and more_. The dataset consists of - 10,052 subjects - 12,765 sessions - 49,193 scans. _For further details about the dataset, including the preprocessing, please consult the [dataset paper preprint](https://arxiv.org/pdf/2506.14432)_. ## The FOMO-MRI Dataset Collection FOMO50K is one of four related datasets in the [FOMO-MRI collection](https://huggingface.co/collections/FOMO-MRI/fomo-mri-datasets). **[FOMO300K](https://huggingface.co/datasets/FOMO-MRI/FOMO300K) is the full superset; every other dataset in the collection — including this one — is a subset of it.** The variants exist to offer different trade-offs between size, access requirements, and preprocessing. | Dataset | Scans | Access | Description | |---|---|---|---| | [FOMO300K](https://huggingface.co/datasets/FOMO-MRI/FOMO300K) | 306,207 | Gated (auto-approved) | The full superset across the collection. Gated because some constituent datasets require Data Use Agreements. | | [FOMO260K](https://huggingface.co/datasets/FOMO-MRI/FOMO260K) | 260,927 | Open (CC-BY-NC-SA) | The freely accessible subset of FOMO300K. No login or access request required. | | [FOMO50K](https://huggingface.co/datasets/FOMO-MRI/FOMO50K) *(this dataset)* | 49,193 | Gated (auto-approved) | A co-registered, skull-stripped, and/or defaced subset of FOMO300K. | | [FOMO45K](https://huggingface.co/datasets/FOMO-MRI/FOMO45K) | 46,149 | Open (CC-BY-NC-SA) | The freely accessible subset of FOMO50K. No login or access request required. | > ⚠️ **Do not combine datasets from this collection.** Because each dataset is a subset of FOMO300K, combining them will result in duplicated scans. ## Format All data is provided as NIfTI-files. The dataset is provided as a collection of datasets, each within its own folder `PTXYZ_DatasetName` (e.g., `PT001_OASIS1`). All data has been standardized and preprocessed (including skull stripped, RAS reoriented, co-registered) with the following format: ``` -- PT001_OASIS1 |-- sub_52 |-- ses_1 |-- t1.nii.gz -- PT002_OASIS2 |-- sub_27 |-- ses_1 |-- t1.nii.gz ``` Sessions with multiple scans of the same sequence are named `sequence_x.nii.gz`. Sessions where the sequence information was not available are named `scan_x.nii.gz`. The dataset has been collected from the following public sources: OASIS1, OASIS2, BraTS-GEN, MSD-BrainTumor, IXI, NKI, SOOP, NIMH, DLBS, IDEAS, ARC, MBSR, UCLA, QTAB, AOMIC ID1000. ## Metadata Files The dataset includes the following metadata files, available both in the main folder and in each individual dataset folder: - `participants.tsv`: Contains demographic and clinical information including age, gender, handedness, and subject group (e.g., control, specific diagnosis) - `mapping.tsv`: Links the files in FOMO50K to the original data source scans - `mri_info.tsv`: Includes MRI acquisition information Additionally, the main folder contains: - `FOMO50K_300K_mapping.tsv`: Maps the files that are present in both the 50K and 300K versions of the dataset ## Citation Users must cite the following paper when using the FOMO50K dataset: ```bibtex @article{Cerri2026large, title={A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning}, author={Cerri, Stefano and Munk, Asbj{\o}rn and Llambias, Sebastian N{\o}rgaard and Ambsdorf, Jakob and Machnio, Julia and Nersesjan, Vardan and Hedeager Krag, Christian and Liu, Peirong and Rocamora Garc{\'\i}a, Pablo and Mehdipour Ghazi, Mostafa and Boesen, Mikael and Benros, Michael Eriksen and Iglesias, Juan Eugenio and Nielsen, Mads}, journal={arXiv preprint arXiv:2506.14432}, year={2026}, url={https://arxiv.org/abs/2506.14432} } ``` In addition, users must comply with all attribution requirements of the constituent datasets included in FOMO50K, as specified in the Usage Notes of the paper.

提供机构:
maas
创建时间:
2026-01-22
二维码
社区交流群
二维码
科研交流群
商业服务