Omni-MATH
收藏资源简介:
Omni-MATH是由北京大学等机构创建的一个专为评估大型语言模型(LLMs)在奥林匹克级别数学推理能力上的综合性基准数据集。该数据集包含4428个竞赛级别的数学问题,这些问题被精心分类为超过33个子领域和10个不同的难度级别。数据集的创建过程包括从全球数学竞赛中收集数据,并通过人工注释进行验证,确保数据的高质量和多样性。Omni-MATH旨在为LLMs在复杂数学问题解决和推理能力上提供一个具有挑战性的评估平台,特别是在奥林匹克级别的数学问题上。
Omni-MATH is a comprehensive benchmark dataset developed by Peking University and other institutions, tailored to evaluate the Olympiad-level mathematical reasoning abilities of Large Language Models (LLMs). This dataset includes 4,428 competition-grade mathematical problems, which are meticulously categorized into more than 33 sub-disciplines and 10 distinct difficulty levels. The dataset was constructed by collecting data from global mathematics competitions and validated via manual annotation to ensure high data quality and diversity. Omni-MATH aims to provide a challenging evaluation platform for assessing LLMs' complex mathematical problem-solving and reasoning capabilities, with a particular focus on Olympiad-level mathematical problems.




