Brain_MRI_Dataset
收藏资源简介:
该数据集是一个精心整理的集合,包含2D脑部MRI图像、对应的肿瘤分割掩码以及粗粒度的正面描述文本。它专为医学视觉-语言和分割实验而构建,特别适用于需要关联MRI图像、其病变掩码以及可见肿瘤形态文本描述的任务。数据集共包含22,150个图像-掩码-描述三元组,整合自六个公开或本地重建的脑部MRI分割数据源:BRISC2025_segmentation(4,793对)、BraTS16_MMNeuroOnco(5,000对)、Brain_Tumor_Classification_2D_masked(3,166对)、Figshare_from_mat_all(3,064对)、MMNeuroOnco_tumor_14_t1ce_mask(3,064对)和MMNeuroOnco_tumor_19_mask(3,063对)。主索引文件为CSV格式,包含三列:image_path(指向MRI图像的相对路径)、mask_path(指向肿瘤分割掩码的相对路径)和caption(从MM-NeuroOnco元数据的coarse_description字段提取的正面粗粒度描述)。这些描述文本涵盖了肿瘤相关的粗粒度属性,如肿瘤类型、成像模态、近似位置、大小、形状和扩散模式(当可用时)。数据集适用于多种研究实验,包括基于脑部MRI图像的肿瘤分割、使用图像-掩码-描述三元组的视觉-语言对齐、基于医学图像-文本对的DPO或偏好学习、弱监督或提示引导的分割,以及评估视觉编码器是否保留病变相关信息。需要注意的是,描述文本为粗粒度描述,并非完整的放射学报告;掩码来自可用的源标注或转换文件,不同子集的掩码约定可能不同;数据集仅用于研究目的,不得用于临床决策。
This curated dataset comprises 2D brain MRI images, corresponding tumor segmentation masks, and coarse-grained positive descriptive texts. It is specifically developed for medical vision-language and segmentation experiments, and is particularly suited for tasks requiring correlation between MRI images, their lesion masks, and textual descriptions of visible tumor morphology. The dataset contains a total of 22,150 image-mask-caption triplets, integrated from six publicly available or locally reconstructed brain MRI segmentation datasets: BRISC2025_segmentation (4,793 pairs), BraTS16_MMNeuroOnco (5,000 pairs), Brain_Tumor_Classification_2D_masked (3,166 pairs), Figshare_from_mat_all (3,064 pairs), MMNeuroOnco_tumor_14_t1ce_mask (3,064 pairs), and MMNeuroOnco_tumor_19_mask (3,063 pairs). The primary index file is in CSV format, with three columns: image_path (relative path to the MRI image), mask_path (relative path to the tumor segmentation mask), and caption (positive coarse-grained description extracted from the coarse_description field of the MM-NeuroOnco metadata). These descriptive texts cover coarse-grained tumor-related attributes, including tumor type, imaging modality, approximate location, size, shape, and invasion pattern (when available). The dataset supports a wide range of research experiments, such as tumor segmentation based on brain MRI images, vision-language alignment using image-mask-caption triplets, DPO or preference learning based on medical image-text pairs, weakly-supervised or prompt-guided segmentation, and evaluating whether visual encoders retain lesion-related information. It should be noted that the descriptive texts are coarse-grained and not complete radiology reports; masks are sourced from available original annotations or converted files, and mask conventions may differ across subsets; the dataset is for research purposes only and must not be used for clinical decision-making.




