baber/agieval
收藏资源简介:
AGIEval是一个以人为中心的基准测试,专门设计用于评估基础模型在人类认知和问题解决相关任务中的一般能力。该基准测试来源于20个官方、公开和高标准的入学和资格考试,这些考试面向普通人类考生,如大学入学考试(如中国高考和美国SAT)、法学院入学考试、数学竞赛、律师资格考试和国家公务员考试。
AGIEval is a human-centric benchmark specifically designed to evaluate the general capabilities of foundation models in tasks associated with human cognition and problem-solving. This benchmark is sourced from 20 official, public, and high-standard admission and qualification examinations targeting regular human test-takers, including college entrance examinations (such as China's Gaokao and the US SAT), Law School Admission Test (LSAT), mathematics competitions, bar examinations, and national civil service examinations.
数据集概述
数据集名称
AGIEval
数据集描述
AGIEval是一个以人为中心的基准,专门设计来评估基础模型在与人认知和问题解决相关的任务中的通用能力。该基准源自20个官方、公开、高标准的人类考试,包括大学入学考试(如中国高考和美国SAT)、法学院入学考试、数学竞赛、律师资格考试和国家公务员考试。
数据集用途
用于评估基础模型在人类认知和问题解决任务中的表现。
数据集类别
- 问题回答
- 文本生成
许可证
MIT
语言
英语
引用信息
@misc{zhong2023agieval, title={AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models}, author={Wanjun Zhong and Ruixiang Cui and Yiduo Guo and Yaobo Liang and Shuai Lu and Yanlin Wang and Amin Saied and Weizhu Chen and Nan Duan}, year={2023}, eprint={2304.06364}, archivePrefix={arXiv}, primaryClass={cs.CL} }



