AiHazz/wikipedia
收藏资源简介:
该数据集是一个多语言文本数据集,覆盖了超过300种语言和方言(包括但不限于阿布哈兹语、阿塞拜疆语、阿拉伯语、中文、英语等),基于2023年11月1日的版本。数据集设计用于文本生成和填充掩码任务,支持语言建模和掩码语言建模。数据规模从少于1千条到超过1千万条不等,具体取决于语言配置。数据文件以训练分割形式组织,每个语言对应一个独立的配置和路径。数据集采用知识共享署名-相同方式共享3.0许可证(CC-BY-SA-3.0)和GNU自由文档许可证(GFDL)进行许可。
This dataset is a multilingual text dataset covering over 300 languages and dialects (including but not limited to Abkhaz, Azerbaijani, Arabic, Chinese, English, etc.), based on the version from November 1, 2023. It is designed for text generation and fill-mask tasks, supporting language modeling and masked language modeling. The data size ranges from less than 1,000 to over 10 million entries, depending on the language configuration. Data files are organized into training splits, with each language corresponding to an independent configuration and path. The dataset is licensed under the Creative Commons Attribution-ShareAlike 3.0 License (CC-BY-SA-3.0) and the GNU Free Documentation License (GFDL).



