ARB
收藏资源简介:
ARB数据集包含1207条英文指令,涵盖了数学、物理、生物、化学和法律等领域的复杂推理挑战,深入探讨了更深层次的知识。这些问题包括选择题、简答题和开放式回答形式,采用了一种结合代码、人工评估和模型分析的混合评估方法。数据集的发起者引入了一种基于规则的评估方法,使GPT-4能够对中间推理步骤进行打分。
The ARB dataset encompasses 1207 English instructions that cover complex reasoning challenges across disciplines such as mathematics, physics, biology, chemistry, and law, delving into deeper levels of knowledge. The questions include multiple-choice, short answer, and open-ended response formats, employing a mixed evaluation method that integrates coding, human assessment, and model analysis. The initiators of the dataset introduced a rule-based evaluation approach to enable GPT-4 to score intermediate reasoning steps.
Advanced Reasoning Benchmark (ARB) 数据集概述
基本信息
- 名称: Advanced Reasoning Benchmark (ARB)
- 维护机构: DuckAI
- 合作机构: 乔治亚理工学院、苏黎世联邦理工学院、Nomos AI、斯坦福大学法律信息学中心、Mila - Quebec AI Institute
- 许可证: MIT
- 相关论文: arXiv:2307.13692
数据集简介
ARB是一个新颖的基准测试数据集,由高级推理问题组成,旨在评估大型语言模型(LLMs)在文本理解和专业领域推理方面的能力。该数据集比现有基准更具挑战性,包含测试数学、物理、生物、化学和法律领域深层知识的问题。
API访问
- 端点URL: https://advanced-reasoning-benchmark.netlify.app/api/
- 完整REST API文档: API文档




