CMB
收藏资源简介:
CMB是一个全面的中文医学基准数据集,由香港中文大学深圳的研究团队开发。该数据集包含280,839个多选题,涵盖6大类别和28个子类别,旨在评估大型语言模型在医学领域的应用。CMB-Exam部分包含资格考试的多选题,而CMB-Clin则包含复杂的临床诊断问题,源自真实的病例研究。数据集的创建过程包括从公开可用的考试题目和课程练习中收集数据,并由专家提供明确的解决方案。CMB数据集的应用领域广泛,旨在解决医学领域中大型语言模型的评估问题,特别是在中国本土文化和语言框架下的应用。
CMB is a comprehensive Chinese medical benchmark dataset developed by a research team at The Chinese University of Hong Kong, Shenzhen. It contains 280,839 multiple-choice questions spanning 6 major categories and 28 subcategories, designed to evaluate the performance of large language models (LLMs) in the medical field. The CMB-Exam subset consists of multiple-choice questions from medical qualifying examinations, while CMB-Clin includes complex clinical diagnostic questions derived from real-world case studies. The dataset was constructed by collecting data from publicly available exam questions and course exercises, with explicit solutions provided by medical experts. With broad application prospects, the CMB dataset aims to address the evaluation challenges of large language models in the medical domain, particularly for applications within the framework of Chinese local culture and language.

- 1CMB: A Comprehensive Medical Benchmark in Chinese香港中文大学深圳 · 2024年



