AI Crawler Directives in the robots.txt of Popular Websites (2026)
收藏官方服务:
资源简介:
A reproducible point-in-time snapshot (2026-07-21) of whether 20 widely used websites reference major AI and LLM crawlers (OpenAI GPTBot, Anthropic ClaudeBot, Google-Extended, Common Crawl CCBot, PerplexityBot, ByteDance Bytespider) in their robots.txt. Derived from each public robots.txt file; fully reproducible; no fabricated values. A "1" records a directive naming the crawler (in practice usually a restriction), not a verified Disallow. Compiled by alexi.sh. Related coverage of AI, privacy and the shift to answer engines: alexi.sh.
提供机构:
Zenodo创建时间:
2026-07-21



