Nemotron-Pretraining-SFT_translated
收藏资源简介:
本数据集是Nemotron Pretraining SFT数据的斯洛文尼亚语翻译版本,源数据取自nvidia/Nemotron-Pretraining-SFT-v1数据集的Nemotron-SFT-General子集,并进行了子采样。使用cjvt/GaMS-DPO-Translator模型进行自动翻译,包含约90万个文档,翻译后的斯洛文尼亚语文本约含8.89亿单词。每个样本包含以下字段:id(保留原始标识符)、text(原始英文文本)、metadata(保留原始元数据,包括模型和类别信息)、chopped_document(为适应翻译模型上下文窗口而分割的较小文本单元列表)、text_sl(英文文本的斯洛文尼亚语翻译)。该数据集适用于斯洛文尼亚语的自然语言处理任务,特别是机器翻译、多语言语言模型预训练与指令微调、以及低资源语言的大语言模型研究。在PoVeJMo研究计划(自适应自然语言处理与大语言模型)和SloLLaMai研究项目(斯洛文尼亚语高效开源模型)支持下开发,采用CC BY 4.0许可证发布。
This dataset is a Slovenian translation version of the Nemotron Pretraining SFT data, sourced from the nvidia/Nemotron-Pretraining-SFT-v1 dataset, specifically a subsample of the Nemotron-SFT-General subset. It was automatically translated using the cjvt/GaMS-DPO-Translator model. The dataset contains approximately 900,000 documents, with the translated Slovenian text comprising about 889 million words. Each sample includes the following fields: id (retaining the original dataset identifier), text (original English text), metadata (retaining original metadata, including model and category information), chopped_document (a list of smaller text units segmented to fit the translation models context window), and text_sl (Slovenian translation of the English text). The dataset is suitable for Slovenian natural language processing tasks, particularly machine translation, multilingual language model pretraining and instruction fine-tuning, and large language model research for low-resource languages. It was developed with support from the PoVeJMo research initiative (Adaptive Natural Language Processing and Large Language Models) and the SloLLaMai research project (Efficient Open Models for Slovenian), and is released under the CC BY 4.0 license.
数据集概述
- 数据集名称:Nemotron Pretraining SFT Slovene translation
- 来源:基于 nvidia/Nemotron-Pretraining-SFT-v1 数据集,选取其子集 Nemotron-SFT-General 进行斯洛文尼亚语翻译。
- 翻译模型:使用 cjvt/GaMS-DPO-Translator 模型完成翻译。
数据规模
- 文档数量:约 900,000 个文档(训练集具体包含 907,416 个样本)
- 斯洛文尼亚语词汇量:翻译后的文档包含约 8.89 亿个斯洛文尼亚语单词
- 数据集大小:
- 下载大小:6,407,592,816 字节
- 数据集大小:14,592,578,522 字节
数据结构
数据集包含以下字段:
id:原始数据集中的唯一标识符(字符串类型)text:原始英文示例(字符串类型)metadata:原始数据集的元数据,包含:models_used:使用的模型(字符串类型)category:类别(字符串类型)
chopped_document:为适应翻译模型上下文窗口而切分的较小子文本单元列表(字符串序列)text_sl:英文示例的斯洛文尼亚语翻译(字符串类型)
许可协议
数据集采用 CC BY 4.0 许可证发布。
相关项目与致谢
- 数据集开发基于 PoVeJMo 研究计划(自适应大语言模型自然语言处理),特别是其中的子项目 SloLLaMai(面向斯洛文尼亚语的开源高效计算模型)。
- 该研究计划由斯洛文尼亚研究与创新机构(ARIS)和 NextGenerationEU 通过复苏与韧性计划资助。
- 作者同时感谢斯洛文尼亚研究与创新机构的核心研究资助(编号 P6-0411 — 斯洛文尼亚语语言资源与技术)。
- 本项目还获得欧盟 Horizon Europe 计划(项目编号 101186647 – AI4DH)资助。
引用信息
bibtex misc{vreš2026buildingstronginstructionlanguage, title={Building a Strong Instruction Language Model for a Less-Resourced Language}, author={Domen Vreš and Tjaša Arčon and Timotej Petrič and Dario Vajda and Marko Robnik-Šikonja and Iztok Lebar Bajec}, year={2026}, eprint={2603.01691}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2603.01691}, }




