open-index/cc-host-dataset
收藏资源简介:
CC Host Dataset — CC-MAIN-2026-21是一个基于Common Crawl的URL索引数据集,提供主机级别的排名信号。该数据集覆盖整个Common Crawl,每个被爬取的URL都作为一行数据,包含20个原始CDX字段和来自CC网络图的谐波中心性排名。数据集大约包含19亿个URL,涉及约2.62亿个主机。它支持多种用途,如URL查询、内容变化检测和多爬取比较,并采用Parquet格式存储,适用于特征提取和文本分类任务。数据来源于Common Crawl,遵循ODC-By许可证。
CC Host Dataset — CC-MAIN-2026-21 is a per-URL index of the entire Common Crawl with host-level rank signals. Every URL captured by the crawler becomes one row, enriched with 20 raw CDX fields and harmonic centrality rank from the CC web graph. The dataset covers roughly 1.9 billion URLs across ~262 million hosts. It supports various use cases such as URL lookup, content change detection, and multi-crawl comparison, stored in Parquet format, and is suitable for tasks like feature extraction and text classification. The data is derived from Common Crawl and released under the ODC-By license.




