GPBench
收藏资源简介:
GPBench是一个基于中山大学附属第六医院和广东人工智能与数字经济实验室(深圳)的通用实践者能力模型构建的中文通用实践者基准测试数据集。该数据集旨在评估大型语言模型在临床诊断和治疗支持方面的能力,并帮助识别模型的理论积累不足。它由三个测试集组成:MCQ测试集、临床案例测试集和AI患者测试,均由专家详细注释,作为评估和分析和当前最先进的大型语言模型的基础。
GPBench is a Chinese general practitioner benchmark dataset developed for building general practitioner capability models, based on the Sixth Affiliated Hospital of Sun Yat-sen University and Guangdong Laboratory of Artificial Intelligence and Digital Economy (Shenzhen). This dataset aims to evaluate the capabilities of large language models in clinical diagnosis and treatment support, and help identify insufficient theoretical accumulation of these models. It consists of three test sets: the MCQ test set, the clinical case test set, and the AI patient test, all of which are meticulously annotated by experts and serve as the basis for evaluating and analyzing current state-of-the-art large language models.

- 1GPBench: A Comprehensive and Fine-Grained Benchmark for Evaluating Large Language Models as General Practitioners中山大学附属第六医院, 广东人工智能与数字经济实验室(深圳), 新兴人民医院, 中山大学智能系统工程学院 · 2025年



