finetranslations-filtered
收藏资源简介:
该数据集名为 'finetranslations-filtered',是一个多语言文本数据集,包含来自 'finetranslations' 的过滤句子,移除了大多数非英语语言中的英文文本。数据集支持多种语言,包括但不限于非洲语、阿姆哈拉语、阿拉伯语、孟加拉语、中文等。数据以 Parquet 格式存储,每个文件包含一个 'sentence' 列。数据集适用于文本分类、语言建模和掩码语言建模等任务。数据集规模在 1M 到 10M 之间,仅包含训练集。数据集采用 Open Data Commons Attribution License (ODC-By) v1.0 许可,使用时需遵守 CommonCrawl 的使用条款。
This dataset, named 'finetranslations-filtered', is a multilingual text dataset containing filtered sentences sourced from 'finetranslations', with English text within most non-English language contents removed. The dataset supports multiple languages, including but not limited to African languages, Amharic, Arabic, Bengali, Chinese, and more. The data is stored in Parquet format, with each file containing a 'sentence' column. This dataset is suitable for tasks such as text classification, language modeling, and masked language modeling. It has a scale ranging from 1 million to 10 million instances, and only includes a training split. The dataset is licensed under the Open Data Commons Attribution License (ODC-By) v1.0, and users must comply with the terms of use of CommonCrawl when utilizing the dataset.




