MATHHAY
收藏资源简介:
MATHHAY是由Salesforce AI Research和新加坡管理大学共同创建的一个自动化基准数据集,旨在评估大型语言模型在长上下文环境中的数学推理能力。该数据集涵盖了从单一文档到多文档、单一步骤到多步骤的多种难度级别的数学推理任务,适用于32K到128K tokens的长度。MATHHAY的创建过程包括文档收集、问题生成、质量控制和海量文档构建四个主要阶段,确保数据集的高质量和真实性。该数据集主要应用于评估和提升大型语言模型在实际场景中的数学推理能力,特别是在需要处理大量文本和复杂数学计算的领域。
MATHHAY is an automated benchmark dataset co-developed by Salesforce AI Research and Singapore Management University, designed to evaluate the mathematical reasoning capabilities of large language models (LLMs) in long-context scenarios. This dataset encompasses mathematical reasoning tasks across diverse difficulty levels, ranging from single-document to multi-document configurations, and single-step to multi-step procedures, with context lengths spanning from 32K to 128K tokens. The construction of MATHHAY consists of four core stages: document collection, question generation, quality control, and large-scale document construction, which ensures the high quality and authenticity of the dataset. This dataset is primarily applied to assess and enhance the mathematical reasoning abilities of LLMs in real-world scenarios, especially in domains requiring processing massive text and executing complex mathematical computations.




