danish-word-problems-v1
收藏资源简介:
本数据集(danish-word-problems,v1 版本)是一个用于训练模型解决丹麦语数学应用题(Word Problems)的合成数据集。该数据集由基础程序生成器构建,其数据生成配方包含“混合”(Mixture)方法。它主要用于对大型语言模型进行监督微调(SFT),旨在提升模型解决丹麦语数学文字问题的能力。该数据集的性能通过在 GSM8K[da](丹麦语版)基准测试上的表现进行评估,使用 v1 数据训练的模型在该测试中取得了 4.86% 的准确率。需要注意的是,此 v1 版本已被更新的 v2 版本所取代。v2 版本在 v1 的基础上进行了显著扩展和改进,包括引入了 16 种组合配方(相较于 v1 的 8 种基础类型)、增加了约 30% 的反向推理链以训练方程逆推能力、为每个数学操作符提供了约 10-15 个模板变体以增加表面多样性,并保留了“混合”配方。建议所有新的监督微调任务均使用 v2 版本。
This dataset (danish-word-problems, version v1) is a synthetic dataset designed for training models to solve Danish-language mathematical word problems. It is constructed using a basic program generator with a data generation recipe that includes the Mixture method. The dataset is primarily used for supervised fine-tuning (SFT) of large language models, aiming to enhance their ability to solve Danish mathematical text problems. Its performance is evaluated on the GSM8K[da] (Danish version) benchmark, where models trained with v1 data achieved an accuracy of 4.86%. Note that this v1 version has been superseded by the newer v2 version. The v2 version introduces significant expansions and improvements over v1, including 16 combination recipes (compared to 8 basic types in v1), an increase of approximately 30% in reverse reasoning chains to train equation inversion capabilities, about 10-15 template variations per mathematical operator to enhance surface diversity, and retention of the Mixture recipe. It is recommended that all new supervised fine-tuning tasks use the v2 version.
- 数据集名称:danish-word-problems-v1
- 状态:已弃用(DEPRECATED)
- 替代数据集:jensjepsen/danish-word-problems-v2
- 替代版本新增内容:
- 16种组合式配方(wp_compose),原版本为8种基础类型
- 反向链(约30%的行),用于训练方程反转
- 每个运算符的习语库(每个运算符约10-15种模板变体),增加表面多样性
- 保留了混合作为同级配方
- 构建方式:仅基于基础程序生成器构建
- 评估结果:使用v1训练的SFT模型在GSM8K[da]评估中得分为4.86%;v2版本解决了在esperanto-word-problems-v4上的同等差距




