ThatDeveloperGuy13/crawler-signal-reference
收藏资源简介:
Crawler Signal Reference Library是一个包含30个文档的参考库,覆盖了网络爬虫从网站读取的所有HTML和HTTP信号。该数据集由Joseph W. Anady(ThatDevPro和ThatDeveloperGuy的创始人)创建。它分为两部分:HTML信号(18个文档),包括应用和渐进式Web应用元数据、规范链接和替代链接、网站图标、所有meta标签(如作者、字符集、颜色方案、内容语言、版权、生成器、关键词、引用者、刷新、机器人、主题颜色、视口)、Open Graph协议、Twitter Cards以及搜索引擎验证;HTTP信号(12个文档),包括状态码(2xx、3xx、4xx、5xx)和头部信息(如缓存、内容、CORS、性能、速率控制、请求、安全、SEO)。数据集旨在为文本检索和文本分类任务提供参考,适用于网络爬虫、搜索引擎优化(SEO)和人工智能搜索等领域。许可证为CC BY 4.0,要求署名。
Crawler Signal Reference Library is a 30-document reference library covering every HTML and HTTP signal a web crawler reads from a site. Authored by Joseph W. Anady, founder of ThatDevPro and ThatDeveloperGuy. It includes HTML signals (18 documents) such as app and PWA metadata, canonical and alternate links, favicon and icons, all meta tags (e.g., author, charset, color-scheme, content-language, copyright, generator, keywords, referrer, refresh, robots, theme-color, viewport), Open Graph protocol, Twitter Cards, and search engine verification. HTTP signals (12 documents) cover status codes (2xx, 3xx, 4xx, 5xx) and headers (e.g., caching, content, CORS, performance, rate-control, request, security, SEO). The dataset is designed for text retrieval and text classification tasks, applicable in web crawling, SEO, and AI search. License: CC BY 4.0 with attribution required.



