CORTEX (Clinically Organized Reasoning and sTructured EXplanation)
收藏资源简介:
CORTEX是由穆罕默德·本·扎耶德人工智能大学、哈索·普拉特纳研究所及哈利法大学联合构建的胸部CT三维结构化推理基准数据集,旨在为医学多模态大语言模型提供可解释的临床推理监督数据。该数据集基于公开的大规模胸部CT数据集CT-RATE构建,包含76,177条经过验证的四阶段推理轨迹,涵盖开放式视觉问答、封闭式视觉问答及报告生成三类任务,每条轨迹均包含任务理解、视觉观察、诊断推理和答案合成的结构化步骤。数据生成过程采用前沿大语言模型结合临床医生设计的提示模板进行多温度采样,并通过基于规则的过滤和五维评分准则进行质量验证,确保了推理的临床准确性与逻辑一致性。该数据集主要应用于三维胸部CT的可信推理模型训练与评估,致力于解决医学影像诊断中因缺乏结构化监督与分阶段验证协议而导致模型推理不可追溯、难以验证的核心问题。
CORTEX is a 3D structured reasoning benchmark dataset for chest CT, jointly constructed by Mohammed bin Zayed University of Artificial Intelligence (MBZUAI), Hasso Plattner Institute (HPI), and Khalifa University. It aims to provide interpretable clinical reasoning supervision data for medical multimodal large language models (LLMs). Built upon the publicly available large-scale chest CT dataset CT-RATE, this dataset contains 76,177 validated four-stage reasoning trajectories covering three types of tasks: open-ended visual question answering (VQA), closed-ended visual question answering, and report generation. Each trajectory includes structured steps of task understanding, visual observation, diagnostic reasoning, and answer synthesis. The data generation process adopts cutting-edge large language models combined with prompt templates designed by clinicians for multi-temperature sampling, and conducts quality validation through rule-based filtering and a five-dimensional scoring criterion, ensuring the clinical accuracy and logical consistency of the reasoning. This dataset is primarily applied to the training and evaluation of trustworthy reasoning models for 3D chest CT, and is committed to addressing the core issues of untraceable and hard-to-verify model reasoning in medical image diagnosis caused by the lack of structured supervision and staged verification protocols.

- 1CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs穆罕默德·本·扎耶德人工智能大学; 哈索·普拉特纳研究所; 哈利法大学 · 2026年



