Dead-Web Common Crawl: a longitudinal host-reachability panel (2018-2026)
收藏资源简介:
# Dead-Web Common Crawl — a longitudinal host-reachability panel (2018-2026)An open dataset labeling the reachability of **172,959,928 registered domains** across **80 monthlyCommon Crawl archives** (2018-2026), built from Common Crawl's robotstxt subset and **calibratedagainst a live re-probe** to separate genuinely-dead domains from those merely blocking crawlers ordropped from the crawl.**~40M domains (~23% of the 8-year union) are calibrated genuinely dead.** Of domains that disappearedfrom Common Crawl, only **64.2% [63.5-64.9%]** are truly dead — 31% still resolve, 4.7% are blocked.Domains last seen *blocked* are only 44.7% dead (the dead-vs-blocked distinction, measured).Per-domain schema: domain, first_ord, last_ord, total_crawls(80), n_present, last_state,recent_absence_streak, label (alive_present/dead_candidate/dark_ambiguous/intermittent).Built with the open-source crawlora-deadweb CLI (https://github.com/Crawlora-org/crawlora-deadweb).Source repo: https://github.com/Crawlora-org/dead-web-commoncrawl . License: CC BY 4.0.



