AI-Crawler Blocking Index
收藏官方服务:
资源简介:
A `robots.txt` census of the **Tranco top 1,000,000 domains** (June 2026): which AI crawlers does each site block? Every domain's `/robots.txt` is fetched and parsed for 20 AI-crawler user-agents (GPTBot, CCBot, ClaudeBot, Google-Extended, Bytespider, …) and whether each is `Disallow: /`'d. This is the dataset behind the Crawlora **AI-Crawler Blocking Index**. CC BY 4.0. Data behind the Crawlora study at https://crawlora.net/ai-crawler-index. Source repository: https://github.com/Crawlora-org/ai-crawler-blocking-index-data.
提供机构:
Zenodo创建时间:
2026-07-06



