dataset-multilingue-750k
收藏资源简介:
Synapse多语言数据集是一个基于OPUS-100构建的高质量多语言平行语料库,专为自然语言处理任务设计。该数据集包含10种语言:葡萄牙语(pt)、英语(en)、西班牙语(es)、法语(fr)、意大利语(it)、德语(de)、阿拉伯语(ar)、丹麦语(da)、捷克语(cs)和希腊语(el)。数据总量为75万个样本,采用JSONL格式,并以Apache-2.0许可证发布。数据集已划分为训练集(675,000个样本)、验证集(37,500个样本)和测试集(37,500个样本),划分过程随机进行以避免数据泄露。每个数据样本包含两个字段:idioma(语言代码)和texto(文本内容)。所有数据均经过清洗、标准化和去重处理,包括移除重复项、规范化空格、清理无效字符以及随机打乱。该数据集适用于多种NLP任务,如语言识别、机器翻译、多语言模型微调与预训练、文本分类以及多语言模型的基准测试。数据集由Synapse BR社区维护,是其开源自然语言处理资源计划的一部分。
The Synapse multilingual dataset is a high-quality multilingual parallel corpus built based on OPUS-100, specifically designed for natural language processing tasks. It includes 10 languages: Portuguese (pt), English (en), Spanish (es), French (fr), Italian (it), German (de), Arabic (ar), Danish (da), Czech (cs), and Greek (el). The total data volume is 750,000 samples, stored in JSONL format, and released under the Apache-2.0 license. The dataset is divided into a training set (675,000 samples), a validation set (37,500 samples), and a test set (37,500 samples), with the division performed randomly to prevent data leakage. Each data sample contains two fields: idioma (language code) and texto (text content). All data has undergone cleaning, standardization, and deduplication, including removal of duplicates, normalization of spaces, cleaning of invalid characters, and random shuffling. This dataset is suitable for various NLP tasks, such as language identification, machine translation, multilingual model fine-tuning and pre-training, text classification, and benchmarking of multilingual models. It is maintained by the Synapse BR community as part of its open-source natural language processing resource initiative.




