Wikipedia Hourly Page View Time Series Dataset (2024)
收藏资源简介:
This dataset provides a comprehensive hourly page view time series for the most frequently visited Wikipedia pages throughout 2024. Scope and Filtering: To ensure data quality and focus on popular and heavy-tailed content, each monthly subset includes only those pages that maintained a minimum threshold of 10 views per day, every day for that month. Scale: This filtering results in a robust dataset covering approximately 3 million unique Wikipedia pages and over 2 billion discrete data points per month. Storage Efficiency: The data is stored in the Apache Parquet columnar format. This allows for high-dimensional time series data, originally exceeding 1.9 billion points, to be compressed into highly manageable files of approximately 1.3 GB per month, optimizing both storage footprint and query performance. File Name: data-202401.parquet is the dataset of January 2024. Citing this dataset? Xinyu Chen, HanQin Cai, Lijun Ding, Jinhua Zhao (2026). TailedTS: Benchmark Dataset for Heavy-Tailed Time Series Prediction and Periodicity Quantification. arXiv:2605.16361. or @misc{chen2026tailedts, title={TailedTS: Benchmark Dataset for Heavy-Tailed Time Series Prediction and Periodicity Quantification}, author={Xinyu Chen and HanQin Cai and Lijun Ding and Jinhua Zhao}, year={2026}, eprint={2605.16361}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2605.16361}, } Explore the Wikipedia page view time series? Source: Long-term Wikipedia page view observations from Analytics Datasets - Pageviews since May 2015.



