cjvt/Nemotron-CC-Math-4plus_translated
收藏资源简介:
该数据集是Nemotron CC-Math数据的斯洛文尼亚语翻译版本,源数据集来自nvidia/Nemotron-CC-Math-v1,是4plus子集的子样本。数据集使用cjvt/GaMS-DPO-Translator模型进行翻译,包含96万份文档,翻译后的文档约有8.74亿斯洛文尼亚语单词。数据结构包括id(从原始数据集保留)、text(原始英文示例)、metadata(从原始数据集保留)、chopped_document(适合翻译模型上下文窗口的较小文本单元列表)和text_sl(英文示例的斯洛文尼亚语翻译)。该数据集在PoVeJMo研究项目(特别是SloLLaMai项目)中开发,受斯洛文尼亚研究机构和欧盟资助,旨在支持斯洛文尼亚语语言模型研究,采用CC BY 4.0许可证。
This dataset is a Slovene translation of the Nemotron CC-Math data, sourced from nvidia/Nemotron-CC-Math-v1 and is a subsample of the 4plus subset. It was translated using the cjvt/GaMS-DPO-Translator model. The dataset contains 960k documents, with translated documents comprising approximately 874 million Slovene words. The data structure includes id (retained from the original dataset), text (original English example), metadata (retained from the original dataset), chopped_document (a list of smaller text units fitting the translator models context window), and text_sl (Slovene translation of the English example). Developed within the PoVeJMo research program, specifically the SloLLaMai project, it is funded by Slovenian research agencies and the European Union to support Slovenian language model research, and is released under the CC BY 4.0 license.




