Reasoning Bench (R-Bench)
收藏资源简介:
R-Bench是一个用于评估语言和多媒体模型推理能力的高水平、多学科、英汉双语的基准数据集。它包含1094个语言模型评价问题和665个多模态模型测试问题,涵盖108个学科。这些问题经过精心挑选,确保难度校准、学科平衡和跨语言对齐,使其成为一项奥林匹克级的多学科基准。该数据集旨在解决现有推理基准在评估复杂推理能力方面的不足,特别是在多学科和多模态情境下。数据集的创建过程包括数据收集、筛选和改进等多个步骤,并通过专家筛选、模型筛选和人工审查进行三次筛选,确保问题的质量。此外,R-Bench还具有多语言特性,通过手动构建选项和翻译,使其能够评估模型在不同语言下的推理能力。
R-Bench is a high-caliber, multi-disciplinary, English-Chinese bilingual benchmark dataset for evaluating the reasoning capabilities of language and multimedia models. It contains 1094 language model evaluation questions and 665 multimodal model test questions, covering 108 disciplines. These questions are carefully selected to ensure difficulty calibration, discipline balance and cross-language alignment, making it an Olympic-grade multi-disciplinary benchmark. This dataset aims to address the limitations of existing reasoning benchmarks in assessing complex reasoning capabilities, especially in multi-disciplinary and multimodal scenarios. The creation process of the dataset includes multiple steps such as data collection, screening and refinement, and undergoes three rounds of screening via expert review, model validation and manual inspection to ensure the quality of the questions. In addition, R-Bench features a multilingual design, which enables the evaluation of models' reasoning capabilities across different languages by manually constructing answer options and performing translations.



