MMRad-IVL-22K
收藏资源简介:
MMRad-IVL-22K是第一个为胸部X光解读中的原生交错视觉语言推理设计的大规模数据集。它反映了放射科医生的重复推理和视觉检查工作流程,包含22K高质量且经过专家验证的多模态诊断轨迹。
MMRad-IVL-22K is the first large-scale dataset specifically designed for native interleaved visual-language reasoning in chest X-ray interpretation. It captures the iterative reasoning and visual inspection workflow of radiologists, and includes 22K high-quality, expert-validated multimodal diagnostic trajectories.
数据集概述
数据集名称
Thinking like a radiologist (MMRad-IVL-22K)
核心描述
MMRad-IVL-22K 是首个为胸部X光片解读中的原生交错视觉语言推理而设计的大规模数据集。它反映了放射科医生重复的推理和视觉检查工作流程,包含 22K 条高质量、经过专家验证的多模态诊断轨迹。
关键特性
- 规模:包含 22,000 条数据。
- 质量:数据为高质量且经过专家验证。
- 模态:多模态诊断轨迹。
- 设计目标:用于胸部X光片解读中的解剖学引导的交错视觉语言推理。
数据状态
根据项目计划,数据集尚未完全发布。
- [ ] 发布 MMRad-IVL 数据集的子集
- [ ] 发布完整的 MMRad-IVL 数据集
相关资源
- 论文:https://arxiv.org/abs/2602.12843
- 代码仓库:https://github.com/qiuzyc/thinking_like_a_radiologist
- 参考模型:
- Anole-7b: https://huggingface.co/GAIR/Anole-7b-v0.1
- Anole-Zebra-CoT: https://huggingface.co/multimodal-reasoning-lab/Anole-Zebra-CoT
引用信息
如果该研究对您有帮助,请考虑引用以下论文:
@article{zhao2026thinking, title={Thinking Like a Radiologist: A Dataset for Anatomy-Guided Interleaved Vision Language Reasoning in Chest X-ray Interpretation}, author={Zhao, Yichen and Peng, Zelin and Yang, Piao and Yang, Xiaokang and Shen, Wei}, journal={arXiv preprint arXiv:2602.12843}, year={2026} }



