RACE
收藏资源简介:
RACE是一个大规模阅读理解数据集,由卡内基梅隆大学语言技术研究所创建。数据集包含27,933篇文章和97,687个问题,这些问题是从中国中学生和高中生的英语考试中收集的,由英语教师设计,涵盖广泛的主题和风格。数据集的创建旨在通过专家设计的问题来评估学生的阅读理解能力,特别是推理能力。RACE的应用领域包括机器阅读理解的研究和评估,旨在解决现有数据集在推理需求和主题覆盖上的不足。
RACE is a large-scale reading comprehension dataset developed by the Language Technologies Institute of Carnegie Mellon University. It comprises 27,933 articles and 97,687 questions collected from English examinations administered to Chinese middle and high school students. These questions are crafted by English teachers and cover a diverse array of topics and writing styles. The dataset was created to assess students' reading comprehension abilities, particularly their reasoning capabilities, using expert-designed questions. Applications of RACE span research and evaluation of machine reading comprehension, aiming to address the shortcomings of existing datasets in terms of reasoning requirements and topic coverage.

- 1Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models中国信息处理实验室 · 2024年



