Website Crawler Benchmark 2026: 4 Controlled Runs
收藏资源简介:
Open data behind Website Crawler Benchmark 2026: 4 Controlled Runs. The dataset records four controlled crawl attempts across two public documentation hosts using a five-page ceiling. Files are provided as CSV, JSON, and JSONL with methodology, limitations, and observed results. Canonical methodology and analysis: Website Crawler Benchmark 2026. Source repository: website-crawler-benchmark-2026-data on GitHub. This is a reproducible snapshot, not a universal crawler ranking. The sample contains four attempts, two hosts, one operator environment, one date, and one page ceiling. Two Firecrawl attempts timed out locally after 90 seconds; final upstream output quality and cost are unknown. The author operates the compared Apify Actor, so raw rows, unsuccessful attempts, and limitations are included for independent audit.



