CLIR-Bench
收藏资源简介:
CLIR-Bench是一个证据可审计的基准数据集,专门设计用于评估模型在不规则临床时间序列上进行多模态问答的能力。该数据集基于MIMIC-IV数据库中的去识别化重症监护室(ICU)记录构建,包含6,600个多项选择题问答实例,涵盖了11个关键的临床变量和干预信号。问题被系统地分为四个核心能力维度:时间理解、推理、预测和决策制定。每个问题都明确关联了时间戳级别的证据以及特定任务的答案推导规则,使得该基准不仅能评估模型的答案准确性,还能全面评估其证据忠实度(模型输出与证据的一致性)、证据依赖性(模型决策对证据的依赖程度)以及因果敏感性(对因果关系的理解)。CLIR-Bench特别关注临床实践中的常见挑战,如数据稀疏、观测异步、数据缺失和长期视野分析,从而为评估模型能否可靠地基于患者特定时间证据进行推理提供了一个现实且具有诊断性的测试平台。需要注意的是,当前公开的数据集版本不包含源自MIMIC-IV的原始时间序列数据及相关证据,访问这些部分需要研究者通过PhysioNet官方平台申请相应的数据使用权限。
CLIR-Bench is an evidence-auditable benchmark dataset specifically designed to evaluate models ability in multimodal question answering on irregular clinical time series. It is constructed from de-identified intensive care unit (ICU) records in the MIMIC-IV database. The dataset contains 6,600 multiple-choice question-answering instances, covering 11 key clinical variables and intervention signals. The questions are systematically organized into four core competency dimensions: temporal understanding, reasoning, prediction, and decision-making. Each question is explicitly associated with timestamp-level evidence and task-specific answer derivation rules, enabling the benchmark to assess not only answer accuracy but also evidence faithfulness (consistency of model outputs with provided evidence), evidence dependence (extent to which model decisions rely on given evidence), and causal sensitivity (understanding of causal relationships). CLIR-Bench focuses on common clinical challenges such as data sparsity, asynchronous observations, missing data, and long-term horizon analysis, providing a realistic and diagnostic testbed for evaluating whether models can reliably identify and reason based on patient-specific temporal evidence. It should be noted that the currently released dataset version does not include the original time series data and related evidence from MIMIC-IV; access to these components requires researchers to apply for appropriate data usage permissions through the official PhysioNet platform.
数据集概述:CLIR-Bench
CLIR-Bench 是一个面向 不规律临床时间序列的多模态问答 的可审计证据基准数据集。
- 数据来源:基于 MIMIC-IV 中脱敏的 ICU 记录构建。
- 规模与内容:包含 6,600 个多项选择问答实例,涵盖 11 个临床变量和干预信号。
- 能力维度:围绕四个核心能力组织——时间理解、推理、预测和决策。
- 评估特性:每个问题均明确关联时间戳级别的证据和任务特定的答案推导规则,支持对答案准确性、证据忠实性、证据依赖性和因果敏感性进行综合评估。
- 应用场景:聚焦稀疏、异步、缺失和长时间跨度的临床观察,为评估模型识别和推理患者特定时间证据的能力提供真实且诊断性的测试平台。
⚠️ 重要说明
- 当前版本 不包含 来自 MIMIC-IV 的时间序列数据和证据。
- 涉及 MIMIC-IV 相关任务的数据访问需要额外的资质认证和 PhysioNet 的批准。
- 研究人员可通过官方网站申请访问:MIMIC-IV
相关论文
- 论文可在此处下载:PDF




