NuminaMath-Enhanced-CoT-JA-50K
收藏资源简介:
NuminaMath Enhanced CoT Dataset (Japanese 50k Subset) 是一个从NuminaMath CoT数据集派生出来的日语数学数据集,旨在通过让大型语言模型反复思考其步骤来加强日语中的推理过程。该数据集包含50,000个样本,每个样本的英文数学问题和解决方案被翻译成日语,并通过模型生成四个日语解决方案。数据集的结构包括原始问题的索引、来源、英文问题和解决方案、日文翻译问题和解决方案、生成的解决方案以及所有四个生成的解决方案。数据集的生成过程涉及使用google/gemma-2-27b-it模型进行多次推理尝试,并通过精确匹配检查来确定最佳解决方案。数据集的使用受到Apache License 2.0和Gemma使用条款的限制。
The NuminaMath Enhanced CoT Dataset (Japanese 50k Subset) is a Japanese mathematical dataset derived from the NuminaMath CoT Dataset, designed to enhance reasoning processes in Japanese by having large language models repeatedly deliberate over their solution steps. This dataset contains 50,000 samples: for each sample, the original English mathematical problem and its corresponding solution are translated into Japanese, and four Japanese-language solutions are generated via models. The dataset structure includes the index of the original problem, its source, the English problem and solution, the Japanese-translated problem and solution, the optimal generated solution, and all four of the initially generated solutions. The dataset generation process involves using the google/gemma-2-27b-it model to perform multiple rounds of reasoning attempts, and identifying the optimal solution through exact match validation. The usage of this dataset is governed by the Apache License 2.0 and the Gemma Terms of Service.




