遇见数据集

UCSC-VLAA/ClinSeek-Bench

收藏
Hugging Face2026-05-20 更新2026-06-14 收录
官方服务:

资源简介:

ClinSeek-Bench是一个用于评估临床推理的评估套件,最初在论文《ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning》中提出。它评估两种设置下的临床推理:Curated Input(模型从源基准提供的证据包中获取答案)和Automated Evidence-Seeking(模型必须使用ClinSeekAgent工具从原始临床数据中检索证据)。该数据集仅发布元数据,不包含原始数据,因为数据来源于受保护的临床数据集(如MIMIC-IV、MIMIC-CXR等),需要用户通过官方渠道获取。数据集包括两个部分:文本任务(1,800个示例,涵盖45个EHR子任务,涉及风险预测和决策制定场景)和多模态任务(989个示例,结合EHR和CXR数据,包括CXR发现存在性、枚举、时间变化比较、24小时恶化预测、住院死亡率预测和表型预测等任务)。元数据文件包括ClinSeek-Bench_text.json和ClinSeek-Bench_multimodal.jsonl,用于本地重建完整基准。

ClinSeek-Bench is an evaluation suite introduced in the paper ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning. It evaluates clinical reasoning under two paired settings: Curated Input, where models answer from evidence packages provided by the source benchmark, and Automated Evidence-Seeking, where models must retrieve evidence from raw clinical data using ClinSeekAgent tools. The dataset releases only metadata, not raw data, as it is built from credentialed clinical datasets (e.g., MIMIC-IV, MIMIC-CXR), which require users to obtain access through official sources. It consists of two splits: text-only tasks (1,800 examples covering 45 EHR subtasks for risk prediction and decision-making scenarios) and multimodal tasks (989 examples combining EHR and CXR data, including tasks such as CXR finding presence, enumeration, temporal change comparison, 24-hour decompensation prediction, in-hospital mortality prediction, and phenotype prediction). The metadata files, ClinSeek-Bench_text.json and ClinSeek-Bench_multimodal.jsonl, are provided to reconstruct the full benchmark locally.

提供机构:
UCSC-VLAA
二维码
社区交流群
二维码
科研交流群
商业服务