cjvt/Nemotron-Pretraining-SFT_translated
收藏资源简介:
该数据集是Nemotron Pretraining SFT数据的斯洛文尼亚语翻译版本,源数据集来自nvidia/Nemotron-Pretraining-SFT-v1,是Nemotron-SFT-General子集的子样本。它包含90万份文档,翻译文档约含8.89亿斯洛文尼亚语单词。数据集结构包括以下字段:id(从原始数据集保留的标识符)、text(原始英文文本示例)、metadata(从原始数据集保留的元数据)、chopped_document(为适应GaMS-DPO-Translator模型上下文窗口而切分的较小文本单元列表)和text_sl(英文示例的斯洛文尼亚语翻译)。翻译工作使用cjvt/GaMS-DPO-Translator模型完成,旨在支持斯洛文尼亚语的自然语言处理任务。
This is a Slovene translation of the Nemotron Pretraining SFT data. The source dataset was taken from nvidia/Nemotron-Pretraining-SFT-v1. It is a subsample of the Nemotron-SFT-General subset. The dataset contains 900k documents, with translated documents containing around 889 million Slovene words. Dataset fields include: id (kept from the original dataset), text (original English example), metadata (kept from the original dataset), chopped_document (the list of smaller text units that fit into the context window of GaMS-DPO-Translator), and text_sl (Slovene translation of the English example). The translation was performed using the cjvt/GaMS-DPO-Translator model.




