遇见数据集

AI-Crawler Blocking Index — robots.txt scan of the Tranco top 1M

收藏
Zenodo2026-06-20 更新2026-06-21 收录
官方服务:

资源简介:

A robots.txt census of the Tranco top 1,000,000 domains (June 2026): which AI crawlers does each site block? Every domain's /robots.txt is fetched and parsed for 20 AI-crawler user-agents (GPTBot, CCBot, ClaudeBot, Google-Extended, Bytespider, …) and whether each is Disallow: /'d. Headline: 998,497 domains scanned; 63.1% serve a robots.txt; 9.33% fully block at least one major AI crawler (14.8% of robots-serving). Most-blocked: GPTBot 7.4% ≈ CCBot 7.2%, then Bytespider 6.8%, ClaudeBot 6.7%, Amazonbot 6.6%. Blocking is concentrated at the head (~13% of the top 10k vs ~9% in the long tail) and among content/publisher sites (News & media 80.6%) far more than transactional ones (e-commerce 11%, search 2.9%). Files: results.jsonl.gz (998,497 rows), a 200-row sample, and headline aggregates. One JSON record per line (domain, rank, scheme, robots_status, has_robots, names_ai, blocks_ai, blocks_any_ai, star_disallow_root). Methodology: one GET /robots.txt per domain from a datacenter IP, parsed into User-agent groups. robots.txt is a published request, not an enforced wall — this measures stated policy, not traffic. Same domain universe as the Anti-Bot Adoption Index, so the two datasets join on domain. License: CC BY 4.0. Source: github.com/Crawlora-org/ai-crawler-blocking-index-data; explorer: crawlora.net/ai-crawler-index.

提供机构:
Zenodo
创建时间:
2026-06-20
二维码
社区交流群
二维码
科研交流群
商业服务