遇见数据集

A Fresh Look at Website Complexity

收藏
Zenodo2026-09-25 更新2026-10-01 收录
官方服务:

资源简介:

Abstract: Over the last 15 years, the web has changed significantly with the use of mobile-first design, JavaScript-heavy single-page applications, HTTP/2 and HTTP/3, and widespread CDN-driven infrastructure. In this replication track paper, we provide a fresh perspective on website complexity based on the current web and revisit the widely cited IMC’11 paper by Butkiewicz et al. [14]. We find non-origin resources becoming increasingly dominant since 2011, now exceeding same-origin content in importance and contributing a substantial share of total traffic, with JavaScript alone accounting for 77.9% of non-origin bytes and Tag Managers appearing on 67.1% of websites. We also analyze which metrics are most critical for predicting page load behavior and find that the number of objects is still the most important factor, followed by total bytes and origins, although correlations have weakened over time. We find that subpages are generally less complex than landing pages, while client-side filtering can significantly reduce the complexity, and mobile and desktop versions show similar complexity. Modern infrastructure is widely adopted, with HTTP/2 used in 58.8% and HTTP/3 in 32.7% of requests, while minification still yields limited benefits for many sites. Modern frontend frameworks are used by around 30% of websites, jQuery remains dominant at 66.8%, and most JavaScript-heavy pages (77.2%) grow after execution, with a median increase of 31.9%. Dataset: All ClickHouse table exports are provided in browser-crawler.zip/web_complexity.zip/domain_categories.zip in the structure <database>/<table>.parquet. We recommend importing into a local ClickHouse to run the notebooks seemlessly. The Jupyter notebooks with data aggregation SQL and plotting code are in plotting.zip.

提供机构:
Zenodo
创建时间:
2026-09-25
二维码
社区交流群
二维码
科研交流群
商业服务