SVAMP_de
收藏资源简介:
SVAMP_de是一个高质量的德语翻译数学数据集,源自SVAMP数据集。它通过先进的LLM技术和严格的验证流程(包括翻译、验证、纠正和结果确认)确保所有数字和变量与英文原版完全一致。数据集专门针对数学应用题,强调在翻译过程中保持数学逻辑和特定数字不变。数据集结构在原始基础上增加了德语字段,如Body_DE、Question_DE和question_concat_DE。
SVAMP_de is a high-quality German-translated mathematical dataset derived from the original SVAMP dataset. It leverages state-of-the-art LLM technologies and a rigorous validation workflow encompassing translation, verification, correction and result confirmation to guarantee full consistency between all numbers and variables and the original English version. This dataset is specifically tailored for mathematical word problems, with a core emphasis on preserving mathematical logic and specific numerical values throughout the translation process. Its structure retains the original framework while adding dedicated German-language fields including Body_DE, Question_DE and question_concat_DE.
SVAMP_de数据集概述
数据集基本信息
- 名称: SVAMP_de
- 语言: 德语、英语
- 许可证: MIT
- 任务类别: 文本生成、问答
- 标签: 数学、应用题、翻译、SVAMP
- 源数据集: ChilleD/SVAMP
数据集描述
SVAMP_de是SVAMP数据集的高质量德语翻译版本。该数据集采用严格的验证流程和最先进的LLM,确保了100%的数值一致性和逻辑保真度。
创建方法
数据生成采用“严格逻辑”流程:
- 翻译: 使用最先进的LLM进行翻译。
- 验证: 通过程序检查每一行数据,确保所有数字与原始英文源数据匹配。
- 修正: 对未通过验证的数据行,使用更严格的提示词重试或手动修补。
- 结果: 在数字和变量上实现了100%的验证通过率。
翻译提示词旨在将语言翻译与数学逻辑解耦,明确指示模型将数字和逻辑视为不可变的“常量”,以防止LLM在翻译过程中尝试“解答”或“本地化”数学问题。
数据结构
与原始数据集结构相同,并增加了德语字段:
Body_DE: 德语正文Question_DE: 德语问题question_concat_DE: 拼接的德语文本
引用信息
bib @misc{vadim_borisov_2025, author = { Vadim Borisov and Richard H. Schreiber }, title = { SVAMP_de (Revision f2bb56b) }, year = 2025, url = { https://huggingface.co/datasets/tabularisai/SVAMP_de }, doi = { 10.57967/hf/7178 }, publisher = { Hugging Face } }




