MR-BEN
收藏资源简介:
MR-BEN是由香港中文大学等机构创建的综合性元推理基准,包含5975个问题,覆盖数学、生物、物理、编程和逻辑等多个学科。数据集通过挑战模型定位和分析自动生成推理步骤中的潜在错误,来评估模型的元推理技能。创建过程中,专家精心设计问题并进行标注,确保数据质量。MR-BEN的应用领域广泛,旨在解决现有评估方法在推理能力评估上的不足,推动AI推理框架的发展。
MR-BEN is a comprehensive meta-reasoning benchmark created by institutions including The Chinese University of Hong Kong and other relevant research organizations. It consists of 5,975 questions spanning multiple disciplines such as mathematics, biology, physics, programming, and logic. This benchmark evaluates the meta-reasoning skills of AI models by challenging them to locate and analyze potential errors in automatically generated reasoning steps. During the development of MR-BEN, experts meticulously designed and annotated all questions to ensure high data quality. With broad application scenarios, MR-BEN aims to address the shortcomings of existing evaluation methods for reasoning ability and promote the advancement of AI reasoning frameworks.

- 1MR-BEN: A Comprehensive Meta-Reasoning Benchmark for Large Language Models香港中文大学 · 2024年



