Web Accessibility Quality Prediction Dataset: HTTP Archive 2022–2025
收藏资源简介:
Structural, performance, and accessibility features for 198,105 web pages extracted from the HTTP Archive (httparchive.crawl.pages) via Google BigQuery across four annual crawls (June 2022–2025). The dataset supports the paper "Predicting and Benchmarking Web Accessibility Quality Across Industry Sectors: A Large-Scale Empirical Study Using the HTTP Archive and Machine Learning." Includes 35 features per page organized into two independent groups: 17 structural/performance features and 18 accessibility features. Also includes BigQuery SQL extraction queries and a complete Python analysis pipeline implementing WAQI construction, dual-experiment ML classification (Random Forest, XGBoost, LightGBM, MLP), SHAP explainability, and statistical tests.



