CUPCase
收藏资源简介:
CUPCase数据集是基于BMC医学案例报告期刊的3562个真实世界案例报告构建的,旨在评估大型语言模型在医学知识提取、诊断、总结等方面的能力。该数据集包含以开放文本格式呈现的诊断和多项选择题形式的案例,涵盖了从oncology(肿瘤学)到obstetrics and gynecology(妇产科)等多种医学学科。数据集的构建过程包括从案例报告中提取案例介绍,移除关于最终诊断的明确提及,并将诊断转化为向量形式以便于模型学习。
The CUPCase dataset is constructed based on 3,562 real-world case reports from the BMC Medical Case Reports journal, and is designed to evaluate the capabilities of large language models (LLMs) in tasks such as medical knowledge extraction, diagnosis, and text summarization. The dataset includes cases presented in both open-text format and multiple-choice question (MCQ) format, covering a wide range of medical disciplines spanning from oncology to obstetrics and gynecology. The dataset construction pipeline involves extracting case introductions from the original case reports, removing explicit references to the final diagnosis, and converting diagnostic information into vector embeddings to facilitate model learning.




