vi-gsm8k-agentic
收藏资源简介:
Vietnamese Grade-School Math (Agentic Self-Instruct) 是一个高质量、原生越南语的小学数学应用题指令微调数据集,旨在解决现有越南语数学数据多依赖机器翻译、质量参差不齐的问题。该数据集采用一种受 Meta Autodata 启发的四子代理自指令管道生成,确保了内容的原生性和逻辑正确性。每个数据样本包含一个越南语数学问题、一个分步的链式思考推理过程、一个经过代码执行验证的最终数值答案,以及生成过程的元数据(如验证状态、难度评估、主题和来源种子)。数据集经过严格清洗,从1500个原始样本中最终保留了1465个高质量样本。实验表明,使用该数据集微调的模型在越南语GSM8K测试集(分布内)和SVAMP衍生测试集(分布外)上的表现均优于使用同等规模机器翻译数据训练的基线模型,尤其在分布外泛化能力上优势明显。该数据集适用于数学推理、指令跟随、链式思考生成等文本生成任务,并为越南语NLP研究和教育应用提供了宝贵资源。
Vietnamese Grade-School Math (Agentic Self-Instruct) is a high-quality, native Vietnamese instruction-tuning dataset for elementary school math word problems. Its core goal is to address the issue of existing Vietnamese math data often relying on machine translation and having inconsistent quality. The dataset is generated using a four-sub-agent self-instruct pipeline inspired by Meta Autodata, ensuring native content and logical correctness. Each data sample includes a Vietnamese math problem, a step-by-step chain-of-thought reasoning process, a final numerical answer verified through code execution, and metadata about the generation process (such as verification status, difficulty assessment, topic, and source seed). The dataset undergoes rigorous cleaning, retaining 1465 high-quality samples from an initial 1500 raw samples. Experiments show that models fine-tuned with this dataset outperform baseline models trained on equivalently sized machine-translated data on both the Vietnamese GSM8K test set (in-distribution) and the SVAMP-derived test set (out-of-distribution), with a particularly notable advantage in out-of-distribution generalization. The dataset is suitable for text generation tasks such as mathematical reasoning, instruction following, and chain-of-thought generation, providing a valuable resource for Vietnamese NLP research and educational applications.





