CaseReportBench
收藏资源简介:
CaseReportBench 是一个专家注释的数据集,用于从临床病例报告中提取密集信息。该数据集包含 138 个病例报告,重点关注罕见疾病,特别是先天性代谢异常。数据集是从 PMC 开放获取子集(CC BY-NC, CC BY-NC-SA, CC BY-NC-ND)中获取的。数据集的构建过程包括数据预处理、病例报告选择、专家引导注释和数据集协调。数据集旨在评估大型语言模型 (LLM) 在从病例报告中提取密集信息方面的能力,并支持罕见疾病的诊断和管理。
CaseReportBench is an expert-annotated dataset for extracting dense information from clinical case reports. It contains 138 case reports focused on rare diseases, particularly inborn errors of metabolism. The dataset is sourced from the PMC Open Access Subset under licenses CC BY-NC, CC BY-NC-SA and CC BY-NC-ND. Its construction process includes data preprocessing, case report selection, expert-guided annotation, and dataset curation. This dataset aims to evaluate the capabilities of Large Language Models (LLMs) in extracting dense information from clinical case reports, and support the diagnosis and management of rare diseases.

- 1CaseReportBench: An LLM Benchmark Dataset for Dense Information Extraction in Clinical Case ReportsUniversity of British Columbia, National Institutes of Health · 2025年



