tum-nlp/PluraMath
收藏资源简介:
PluraMath是一个人工策划的多语言数学推理基准,将PolyMath基准扩展到18种代表性不足的语言,涵盖6个语系,从中等资源语言(如印地语、土耳其语)到极端低资源语言(如上索布语和下索布语,母语者少于1.5万)。每个语言包含500个竞赛风格的数学问题,分为四个难度级别(低、中、高、顶级),所有问题均由母语者验证。数据集用于评估大型语言模型在多语言数学推理上的性能,揭示了高资源语言与代表性不足语言之间的持续差距。数据通过三阶段人工循环流程构建:自动翻译、人工验证和质量控制。
PluraMath is a human-curated multilingual mathematical reasoning benchmark that extends the PolyMath benchmark to 18 additional underrepresented languages spanning 6 language families, from mid-resource languages such as Hindi and Turkish down to extreme low-resource languages such as Upper and Lower Sorbian (< 15k L1 speakers). Every language contains 500 competition-style math problems across four difficulty levels (low, medium, high, top), all validated by native speakers. The dataset is used to evaluate LLMs on multilingual mathematical reasoning, highlighting a persistent gap between high-resource and underrepresented languages. It was constructed through a three-stage, human-in-the-loop pipeline: automatic translation, manual verification, and quality control.




