AI knows your brand, not what you do: an 800-domain dataset of AI search citation, AI crawler blocking, and llms.txt adoption
收藏资源简介:
This dataset accompanies the SearchGrade study "AI knows your brand, not what you do". It contains the raw per-domain records, the derived CSV tables, and the full sample definition for an audit of 800 websites drawn from the Tranco ranking (list 5648N, generated 2026-07-31). Each homepage and its domain-root files were audited once with 114 automated checks covering technical SEO, content, answer engine optimization and generative engine optimization. Google's Gemini (gemini-3.5-flash) with Google Search grounding was then asked about each domain by name, on 2026-08-03. For a 150-domain subsample a second, category level question was asked in the same record minutes apart, with the brand never named, giving 123 paired domains in which each domain acts as its own control. Principal findings. In the paired subsample, 84.6% of domains appear in Google AI's grounded sources when asked about by name, against 27.6% when asked about their own category. Not one domain was cited for its category but not its name. Across the full sample, 573 of 702 domains (81.6%) were cited when asked about by name. 23% of 682 domains block at least one named AI crawler in robots.txt, ranging from 4.8% of government and non-profit sites to 61.4% of news and media. 84.7% publish no llms.txt. Limitations. One engine, one model, one date; nothing here measures ChatGPT. The citation rates are floors rather than point estimates: re-querying 30 domains 1.9 hours later left every cited domain cited, while 4 of 13 uncited domains flipped to cited and none flipped the other way. Scope is the homepage plus domain root files, never a site crawl. 14 of the audit's own content and answer engine checks show a writing system gap and are excluded from every pooled percentage; 12 were traced to defects in our own code. Sector labels are generated by a language model. Blocking is not shown to cause lower citation: the association is confounded by sector. Tranco ranks DNS and resolver prominence, not visits, so this is not a sample of the most visited websites.



