AI-Crawler Blocking Index — robots.txt scan of the Tranco top 1M
收藏资源简介:
A robots.txt census of the Tranco top 1,000,000 domains (June 2026): which AI crawlers does each site block? Every domain's /robots.txt is fetched and parsed for 20 AI-crawler user-agents (GPTBot, CCBot, ClaudeBot, Google-Extended, Bytespider, …) and whether each is Disallow: /'d. Headline: 998,497 domains scanned; 63.1% serve a robots.txt; 9.33% fully block at least one major AI crawler (14.8% of robots-serving). Most-blocked: GPTBot 7.4% ≈ CCBot 7.2%, then Bytespider 6.8%, ClaudeBot 6.7%, Amazonbot 6.6%. Blocking is concentrated at the head (~13% of the top 10k vs ~9% in the long tail) and among content/publisher sites (News & media 80.6%) far more than transactional ones (e-commerce 11%, search 2.9%). Files: results.jsonl.gz (998,497 rows), a 200-row sample, and headline aggregates. One JSON record per line (domain, rank, scheme, robots_status, has_robots, names_ai, blocks_ai, blocks_any_ai, star_disallow_root). Methodology: one GET /robots.txt per domain from a datacenter IP, parsed into User-agent groups. robots.txt is a published request, not an enforced wall — this measures stated policy, not traffic. Same domain universe as the Anti-Bot Adoption Index, so the two datasets join on domain. License: CC BY 4.0. Source: github.com/Crawlora-org/ai-crawler-blocking-index-data; explorer: crawlora.net/ai-crawler-index.



