FineWeb-Edu-Ar
收藏资源简介:
FineWeb-Edu-Ar是由沙特数据与人工智能管理局和阿卜杜拉国王科技大学合作创建的机器翻译数据集,旨在支持阿拉伯语小语言模型的训练。该数据集包含189,405,457条记录,总计202亿个阿拉伯语训练的token,是从HuggingFace的FineWeb-Edu数据集中机器翻译而来。数据集的创建过程包括使用NLLB模型进行翻译,并通过滑动窗口技术处理文本以减少成本和碳排放。FineWeb-Edu-Ar主要用于支持阿拉伯语小语言模型的预训练,旨在解决阿拉伯语高质量教育数据稀缺的问题。
FineWeb-Edu-Ar is a machine translation dataset jointly created by the Saudi Authority for Data and Artificial Intelligence and King Abdullah University of Science and Technology (KAUST), aiming to support the training of small Arabic language models. This dataset contains 189,405,457 records, with a total of 20.2 billion Arabic training tokens, and is machine-translated from HuggingFace's FineWeb-Edu dataset. The dataset development process adopts the NLLB model for translation, and applies a sliding window technique to text processing in order to reduce costs and carbon emissions. FineWeb-Edu-Ar is primarily used to support the pre-training of small Arabic language models, with the goal of addressing the scarcity of high-quality educational data for the Arabic language.




