MedGPTEval
收藏资源简介:
MedGPTEval是由上海人工智能实验室与多家医疗机构合作开发的中文医学数据集,包含27个多轮对话和7个病例报告,总计34个案例。数据集涵盖了14种疾病,从系统性疾病到意外伤害,旨在评估大型语言模型在医学领域的专业能力和社交综合能力。创建过程中,由临床专家设计数据集内容,确保与大型语言模型的交互质量。该数据集主要用于评估模型在医学对话和病例报告处理中的表现,解决模型在医学应用中可能产生的安全风险问题,如幻觉(不完全可靠的响应)。
MedGPTEval is a Chinese medical dataset developed by Shanghai AI Laboratory in collaboration with multiple medical institutions. It comprises 27 multi-turn dialogues and 7 case reports, totaling 34 cases. The dataset covers 14 types of diseases, ranging from systemic disorders to accidental injuries, and aims to evaluate the professional competence and comprehensive social capabilities of large language models (LLMs) in the medical field. During the development process, clinical experts designed the dataset content to ensure the quality of interactions with large language models. This dataset is primarily used to assess models' performance in medical dialogue and case report processing, and to address potential safety risks arising from medical applications of LLMs, such as hallucinations (incompletely reliable responses).

- 1MedGPTEval: A Dataset and Benchmark to Evaluate Responses of Large Language Models in Medicine上海人工智能实验室 · 2023年



