megawika-2
收藏资源简介:
MegaWika 2.0是一个多语言和跨语言文本数据集,包含维基百科的结构化视图,最终涵盖50种语言,包括从所有引用的网页源中干净提取的内容。初始版本基于2024年5月1日的维基百科转储,总共包含约7700万篇文章和7100万条抓取的网页引用。英文集合是最大的,包含约1000万篇文章和2400万条抓取的网页引用。数据集以JSON-lines格式呈现,每个块包含最多1000篇文章,每行是一个独立的JSON编码的维基百科文章。除了文章文本外,还提供了文章的结构视图,分为标题、段落、表格和引用。引用包括一个可抓取的网页源的URL,并包含在该网页源中找到的干净提取的内容。数据集还包括详细的统计数据和改进的引用提取过程。
MegaWika 2.0 is a multilingual and cross-lingual text dataset featuring a structured view of Wikipedia, ultimately covering 50 languages with content cleanly extracted from all cited web sources. The initial version is based on the Wikipedia dump dated May 1, 2024, and contains approximately 77 million articles and 71 million crawled web citations in total. The English subset is the largest, comprising roughly 10 million articles and 24 million crawled web citations. The dataset is formatted in JSON-lines, where each chunk contains up to 1000 articles, and each line is an independent JSON-encoded Wikipedia article. In addition to the article text, a structured view of each article is provided, categorized into titles, paragraphs, tables, and citations. Each citation includes the URL of a crawlable web source, alongside the cleanly extracted content found within that source. The dataset also incorporates detailed statistics and an improved citation extraction pipeline.




