遇见数据集

AI Crawler Directives in the robots.txt of Popular Websites (2026)

收藏
Zenodo2026-07-21 更新2026-08-01 收录
官方服务:

资源简介:

A reproducible point-in-time snapshot (2026-07-21) of whether 20 widely used websites reference major AI and LLM crawlers (OpenAI GPTBot, Anthropic ClaudeBot, Google-Extended, Common Crawl CCBot, PerplexityBot, ByteDance Bytespider) in their robots.txt. Derived from each public robots.txt file; fully reproducible; no fabricated values. A "1" records a directive naming the crawler (in practice usually a restriction), not a verified Disallow. Compiled by alexi.sh. Related coverage of AI, privacy and the shift to answer engines: alexi.sh.

提供机构:
Zenodo
创建时间:
2026-07-21
二维码
社区交流群
二维码
科研交流群
商业服务