A Fresh Look at Website Complexity
收藏资源简介:
Abstract: Over the last 15 years, the web has changed significantly with the use of mobile-first design, JavaScript-heavy single-page applications, HTTP/2 and HTTP/3, and widespread CDN-driven infrastructure. In this replication track paper, we provide a fresh perspective on website complexity based on the current web and revisit the widely cited IMC’11 paper by Butkiewicz et al. [14]. We find non-origin resources becoming increasingly dominant since 2011, now exceeding same-origin content in importance and contributing a substantial share of total traffic, with JavaScript alone accounting for 77.9% of non-origin bytes and Tag Managers appearing on 67.1% of websites. We also analyze which metrics are most critical for predicting page load behavior and find that the number of objects is still the most important factor, followed by total bytes and origins, although correlations have weakened over time. We find that subpages are generally less complex than landing pages, while client-side filtering can significantly reduce the complexity, and mobile and desktop versions show similar complexity. Modern infrastructure is widely adopted, with HTTP/2 used in 58.8% and HTTP/3 in 32.7% of requests, while minification still yields limited benefits for many sites. Modern frontend frameworks are used by around 30% of websites, jQuery remains dominant at 66.8%, and most JavaScript-heavy pages (77.2%) grow after execution, with a median increase of 31.9%. Dataset: All ClickHouse table exports are provided in browser-crawler.zip/web_complexity.zip/domain_categories.zip in the structure <database>/<table>.parquet. We recommend importing into a local ClickHouse to run the notebooks seemlessly. The Jupyter notebooks with data aggregation SQL and plotting code are in plotting.zip.



