Hounsfield-CoT
收藏资源简介:
Hounsfield-CoT是由国际关系大学和特伦托大学联合创建的大规模三维医学推理数据集,旨在解决多模态大语言模型在三维医学影像理解中缺乏显式空间推理监督的关键问题。该数据集包含11,200条高质量实例,源自CT-RATE数据集,通过创新的切片式数据合成范式构建,模拟放射科医师的逐层诊断工作流程,将全局临床先验转化为细粒度的切片观察并合成为可解释的思维链。其构建过程采用观察者-合成器双智能体框架,严格遵循序列空间追踪、真实三维空间感知和鉴别排除三大临床原则,最终形成推理链与答案高度解耦的结构化数据。该数据集主要应用于增强三维医学影像的空间推理能力,旨在解决传统二维预训练模型在理解复杂三维解剖结构时出现的空间幻觉和部分容积效应等问题,为三维临床诊断提供透明且可解释的推理基础。
Hounsfield-CoT is a large-scale 3D medical reasoning dataset jointly created by the University of International Relations and the University of Trento, aiming to address the key challenge that multimodal large language models (LLMs) lack explicit spatial reasoning supervision in 3D medical image understanding. This dataset contains 11,200 high-quality instances derived from the CT-RATE dataset, and is constructed via an innovative slice-wise data synthesis paradigm. It simulates the layer-by-layer diagnostic workflow of radiologists, transforming global clinical priors into fine-grained slice-level observations and synthesizing them into interpretable chain-of-thought. Its construction adopts an observer-synthesizer dual-agent framework, strictly following three core clinical principles: sequential spatial tracking, real 3D spatial awareness, and differential exclusion, ultimately yielding structured data where the reasoning chain and the final answer are highly decoupled. This dataset is primarily applied to enhance the spatial reasoning capabilities for 3D medical image understanding, aiming to resolve issues such as spatial hallucinations and partial volume effects that arise when traditional 2D pre-trained models comprehend complex 3D anatomical structures, thereby providing transparent and interpretable reasoning foundations for 3D clinical diagnosis.
数据集概述
该数据集与论文 “Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models” 相关,旨在提升医学多模态大语言模型在三维空间推理方面的能力。
核心目标
- 解决现有多模态大语言模型在三维体数据(3D volumetric data)处理上的瓶颈,尤其是标注不足的问题。
- 通过数据驱动的方法,将全局临床知识转化为细粒度的逐切片指令轨迹(slice-wise instructional trajectories),增强模型的三维空间追踪和鉴别诊断能力,无需进行资源密集的三维专用预训练。
即将发布的内容
该仓库计划陆续发布以下资源:
- 数据合成管线(Data Synthesis Pipeline):用于从体数据标注中生成结构化思维链(Chain-of-Thought, CoT)描述的脚本。
- 训练与微调(Training & Fine-Tuning):高效的LoRA微调方案,用于将二维视觉-语言基线模型与三维临床逻辑对齐。
- 评估协议(Evaluation Protocols):采用“LLM-as-a-judge”的评估脚本,用于解析三维空间坐标输出结果。
当前状态
- 代码库、合成推理数据集以及评估脚本正在整理中,准备完成后将公开发布。

- 1Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models国际关系大学; 特伦托大学 · 2026年




