ytu-ce-cosmos/aime25-tr
收藏资源简介:
该数据集包含2025年美国数学邀请赛(AIME)问题的土耳其语翻译,旨在作为评估大型语言模型(LLMs)在土耳其语中高级数学推理能力的基准。问题使用GPT-5翻译,并经过人工验证和校正。AIME是介于AMC 10/12和USAMO之间的中级考试,问题设计比标准高中数学更难,需要创造性解决问题和对算术、代数、计数、几何、数论和概率的深入理解。数据集中的每个条目代表AIME 2025竞赛中的一个具体问题,包括索引(唯一标识符)、问题(土耳其语的数学问题陈述)和答案(正确的整数解)。该数据集特别适用于:1. 基准测试:在土耳其语中测试LLMs在困难、多步推理任务上的表现,其中记忆不太可能产生正确结果;2. 思维链(CoT)评估:分析模型在非英语语言中生成有效证明步骤的性能。原始问题来源于美国数学协会(MAA)组织的数学竞赛,数据集在合理使用原则下提供用于研究和教育目的。
This dataset contains the Turkish translations of problems from the 2025 American Invitational Mathematics Examination (AIME). It is intended to serve as a benchmark for evaluating the advanced mathematical reasoning capabilities of Large Language Models (LLMs) in the Turkish language. The questions were translated into Turkish using GPT-5 and subsequently manually verified and corrected. The AIME is an intermediate examination between the AMC 10/12 and the USAMO. The problems are designed to be much more difficult than standard high school mathematics, requiring creative problem-solving and deep understanding of arithmetic, algebra, counting, geometry, number theory, and probability. Each entry in the dataset represents a specific problem from the AIME 2025 competition, including the Index (unique identifier), Problem (the full text statement of the mathematical problem in Turkish), and Answer (the correct integer solution). This dataset is particularly useful for: 1. Benchmarking: Testing LLMs on hard, multi-step reasoning tasks in Turkish where memorization is less likely to yield correct results compared to simpler benchmarks; 2. Chain-of-Thought (CoT) Evaluation: Analyzing model performance in generating valid proof steps in a non-English language. The original problems are sourced from the mathematical competitions organized by the Mathematical Association of America (MAA). This dataset is provided for research and educational purposes under fair use principles.




