LLMEval-Med
收藏资源简介:
LLMEval-Med是一个新的基准数据集,涵盖了五个核心医疗领域,包括从真实世界电子健康记录和专家设计的临床场景中创建的2,996个问题。数据集覆盖了医疗知识、语言理解、推理、文本生成和安全伦理五个维度,并进一步细分为27个次级能力指标。LLMEval-Med采用自动评分和人工评分相结合的方法,确保评分的可靠性和实用性,旨在为医疗领域大型语言模型提供一个全面、系统、真实的评估标准。
LLMEval-Med is a novel benchmark dataset covering five core medical domains, with 2,996 questions created from real-world electronic health records and expert-designed clinical scenarios. The dataset encompasses five dimensions: medical knowledge, language understanding, reasoning, text generation, and safety ethics, and is further subdivided into 27 secondary capability metrics. Adopting a combined approach of automatic and manual scoring to ensure the reliability and practicality of evaluation outcomes, LLMEval-Med aims to provide a comprehensive, systematic and realistic evaluation benchmark for large language models in the medical field.




