CliniQ
收藏资源简介:
ClنيQ数据集是由清华大学等机构创建的,旨在评估电子健康记录中的实体检索性能。该数据集利用MIMIC-III数据集中的出院摘要作为电子健康记录语料库,并以ICD疾病代码、手术代码和处方标签作为查询。数据集包含1000份患者笔记,被划分为16550个片段,收集了1246个独特的查询和77206个详细的相关性判断,是先前数据集规模的十倍以上。该数据集可用于单患者检索和多患者检索两种设置,以应对不同的应用场景,如患者图表审查和患者队列选择等。
The ClنيQ dataset was created by institutions including Tsinghua University to evaluate entity retrieval performance in electronic health records (EHRs). This dataset utilizes discharge summaries from the MIMIC-III dataset as its EHR corpus, with ICD disease codes, surgical procedure codes and prescription labels serving as queries. The dataset contains 1000 patient notes split into 16,550 segments, and includes 1,246 unique queries along with 77,206 detailed relevance judgments, with its scale over ten times that of previous datasets. It supports two retrieval settings: single-patient retrieval and multi-patient retrieval, to accommodate diverse application scenarios such as patient chart review and patient cohort selection.

- 1Evaluating Entity Retrieval in Electronic Health Records: a Semantic Gap Perspective清华大学 · 2025年



