RADAR
收藏资源简介:
RADAR数据集是一个用于训练腹部CT诊断通用视觉-语言模型的大规模医学影像数据集。该数据集旨在支持开发能够实现专家级诊断性能的人工智能模型,覆盖广泛的临床任务。数据集包含超过40万次增强腹部CT检查,从中提取了1500万个解剖层面的图像-文本对。数据直接从临床报告中学习,无需人工标注,涵盖了18个解剖结构和146个影像学发现。数据集包含CT图像和对应的临床报告文本,形成视觉-语言对,适用于医学影像诊断、视觉-语言预训练、腹部CT分析等任务。数据经过预处理,包括图像调整大小和解剖结构分割掩码的生成。
The RADAR dataset is a large-scale medical imaging dataset for training general vision-language models for abdominal CT diagnosis. It aims to support the development of AI models with expert-level diagnostic performance, covering a wide range of clinical tasks. The dataset includes over 400,000 enhanced abdominal CT examinations, from which 15 million image-text pairs at the anatomical level are extracted. Data is learned directly from clinical reports without manual annotation, covering 18 anatomical structures and 146 imaging findings. The dataset contains CT images and corresponding clinical report texts, forming vision-language pairs, suitable for tasks such as medical image diagnosis, vision-language pre-training, and abdominal CT analysis. Data is preprocessed, including image resizing and generation of anatomical structure segmentation masks.
- 数据集名称:RADAR(RADAR: An Expert-Level Generalist AI for Abdominal CT Diagnosis)
- 许可证:CC BY-NC-SA 4.0
- 语言:英语
- 标签:视觉语言预训练、医学、CT诊断
- 数据集规模:包含超过40万例增强腹部CT检查,以及1500万个体素级别的图像-文本对
- 数据来源:直接来自临床报告,无需人工标注
- 覆盖范围:涵盖18个解剖结构和146种影像学发现
- 模型性能:在多个中心和多种场景的内部及外部评估中,实现了高诊断性能和稳健的泛化能力。在读者研究中,AI辅助使26名放射科医生的诊断灵敏度提高了约10%
- 预训练模型与支持文件:提供多个预训练模型检查点,包括RADAR预训练、RADAR+从头训练、RADAR+微调等,以及中英文BERT分词器和模型、视觉分支(UNet)检查点。另提供可选的CT图像和解剖掩码数据(用于演示)
- 相关链接:数据集页面地址为
https://huggingface.co/datasets/radar-generalist/RADAR




