Stratified sample of 20,000 Wikipedia articles across 40 languages, 500 per language
收藏资源简介:
This dataset is a stratified random sample of 20,000 articles from 40 Wikipedia language editions (500 drawn uniformly at random from each), as they stood on 2026-05-01, with the metadata of every revision of each article and its talk pages up to that date. The dataset includes per-revision counts of bytes added and removed and of sections added, derived from the wikitext, along with change tags and page move and deletion log entries. It contains no article or talk page text, and editor names and IDs are pseudonymized. The data was generated with this code, exported using the export-wp-article-dataset command. The repository also computes the article trajectory features built on this data. This dataset is part of the project Exploring Wikimedia Communities, Trace Data, Social Systems and Causality. See the included README.md file for details.



