wikipedia-language-snippets-filtered
收藏资源简介:
该数据集包含从多种语言的维基百科文章中提取的清理过的句子片段,这些片段取自维基百科文章的前半部分,并过滤了存根文章。数据集以Parquet文件格式存储,每个文件包含一个'sentence'列。数据集适用于语言建模任务,如文本生成和掩码语言建模。数据来源于维基百科的dump文件,使用'mwparserfromhell'工具进行解析。原始文本内容根据GNU自由文档许可证和知识共享署名-相同方式共享3.0许可证授权。数据集支持多种语言,包括英语、西班牙语、法语、德语、意大利语、葡萄牙语、荷兰语、越南语、土耳其语、拉丁语、印尼语、马来语、南非荷兰语、阿尔巴尼亚语、冰岛语、挪威语、瑞典语、丹麦语、芬兰语、匈牙利语、波兰语、捷克语、罗马尼亚语、俄语、保加利亚语、乌克兰语、塞尔维亚语、白俄罗斯语、哈萨克语、马其顿语、蒙古语、中文、日语、韩语、印地语、乌尔都语、孟加拉语、泰米尔语、泰卢固语、马拉地语、古吉拉特语、卡纳达语、马拉雅拉姆语、旁遮普语、阿萨姆语、奥里亚语、阿拉伯语、波斯语、普什图语、信德语、维吾尔语、希腊语、希伯来语、亚美尼亚语、格鲁吉亚语、阿姆哈拉语、高棉语、老挝语、缅甸语、泰语、僧伽罗语、斯瓦希里语、提格里尼亚语、他加禄语、藏语、迪维希语、巴斯克语、他加禄语。
This dataset contains cleaned sentence fragments extracted from Wikipedia articles across multiple languages. These fragments are sourced from the first half of Wikipedia articles, with stub articles filtered out. The dataset is stored in Parquet file format, where each file includes a 'sentence' column. It is applicable to language modeling tasks such as text generation and masked language modeling. The data is derived from Wikipedia dump files and parsed using the 'mwparserfromhell' tool. The original text content is licensed under the GNU Free Documentation License and the Creative Commons Attribution-ShareAlike 3.0 License. The dataset supports a wide range of languages, including English, Spanish, French, German, Italian, Portuguese, Dutch, Vietnamese, Turkish, Latin, Indonesian, Malay, Afrikaans, Albanian, Icelandic, Norwegian, Swedish, Danish, Finnish, Hungarian, Polish, Czech, Romanian, Russian, Bulgarian, Ukrainian, Serbian, Belarusian, Kazakh, Macedonian, Mongolian, Chinese, Japanese, Korean, Hindi, Urdu, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Assamese, Odia, Arabic, Persian, Pashto, Sindhi, Uyghur, Greek, Hebrew, Armenian, Georgian, Amharic, Khmer, Lao, Burmese, Thai, Sinhala, Swahili, Tigrinya, Tagalog, Tibetan, Dhivehi, and Basque.




