NeuML/wikipedia-20251101
收藏资源简介:
--- annotations_creators: - no-annotation language: - en language_creators: - found license: - cc-by-sa-3.0 - gfdl multilinguality: - monolingual pretty_name: Wikipedia English November 2025 size_categories: - 1M<n<10M source_datasets: [] tags: - pretraining - language modelling - wikipedia - web task_categories: [] task_ids: [] --- # Dataset Card for Wikipedia English November 2025 Dataset created using this [repo](https://huggingface.co/datasets/NeuML/wikipedia) with a [November 2025 Wikipedia snapshot](https://dumps.wikimedia.org/enwiki/20251101/). This repo also has a precomputed pageviews database. This database has the aggregated number of views for each page in Wikipedia. This file is built using the Wikipedia [Pageview complete dumps](https://dumps.wikimedia.org/other/pageview_complete/readme.html)
--- annotations_creators: - 无注释 language: - 英语(en) language_creators: - 现有资源获取 license: - 知识共享署名-相同方式共享3.0协议(CC BY-SA 3.0) - GNU自由文档协议(GFDL) multilinguality: - 单语言 pretty_name: - 2025年11月版英语维基百科 size_categories: - 100万 < 数据量 < 1000万 source_datasets: [] tags: - 预训练 - 语言建模 - 维基百科 - 网络文本 task_categories: [] task_ids: [] --- # 2025年11月版英语维基百科数据集卡片 本数据集基于此[代码仓库](https://huggingface.co/datasets/NeuML/wikipedia)制作,采用了[2025年11月英语维基百科快照](https://dumps.wikimedia.org/enwiki/20251101/)。 该代码仓库同时提供预计算完成的页面访问量数据库。该数据库汇总了维基百科各页面的累计访问量,其构建基于维基百科的[完整页面访问量转储文件](https://dumps.wikimedia.org/other/pageview_complete/readme.html)



