zabr946/Chatbot-Url
收藏资源简介:
这是一个多语言文本数据集,支持超过300种语言和方言,包括常见语言(如英语、中文、西班牙语)和低资源语言(如阿布哈兹语、阿切语等)。数据集按语言代码组织,每个语言配置对应一个训练数据文件路径,格式为20231101.[语言代码]/train-*,表明数据可能基于2023年11月1日的快照。数据集适用于文本生成和填充掩码任务(如语言建模和掩码语言建模),规模从少于1K到超过1M不等,具体取决于语言。许可证为CC BY-SA 3.0和GFDL,允许共享和改编。
This is a multilingual text dataset supporting over 300 languages and dialects, including common languages such as English, Chinese, and Spanish, as well as low-resource languages like Abkhaz, Ache, and others. The dataset is organized by language codes, where each language configuration corresponds to a training data file path in the format of 20231101.[language code]/train-*, indicating that the data may be based on a snapshot from November 1, 2023. This dataset is applicable to text generation and mask filling tasks such as language modeling and masked language modeling. Its scale ranges from less than 1K to over 1M, varying depending on the specific language. The licenses are CC BY-SA 3.0 and GFDL, which permit sharing and adaptation of the dataset.



