llms.txt across Romanian domains cited by AI engines
收藏资源简介:
First release of the episode 4 dataset. Measured: 9 August 2026, one complete run Frame: 87 Romanian-market domains that AI engines cite when answering Romanian-language questions Licence: CC BY 4.0 Findings 56 of 87 domains (64.4%) publish an llms.txt — already the majority practice in this niche 12 files are plugin-generated, averaging 25,070 bytes against 8,803 for the 44 hand-written ones, and carry neither the H1 nor the summary line the llmstxt.org specification asks for 39 files are spec-conformant; 8 contain no curated link at all The single most-cited domain in the frame serves a plugin dump — evidence against llms.txt being what earns citations 2 domains block answer-engine crawlers, 1 of them while publishing an llms.txt; 1 blocks training scrapers only Contents Dataset as CSV, raw scanner output as JSON, and both scripts — collection and derivation — so the measurement can be repeated or disputed. Study page: https://websem.ro/resurse/aeo/studiu-llms-txt-romania



