遇见数据集

How LLMs Source Brand Reputation Across Languages and Markets: A Cross-Market Citation Dataset (2026)

收藏
Zenodo2026-06-24 更新2026-06-28 收录
官方服务:

资源简介:

The citations that grounded large-language-model answers about brands, merged from three Rankfor.AI studies, supporting the paper "How Large Language Models Source Brand Reputation Across Languages and Markets." It covers 167,551 URL-grounded citations (189,974 total attribution rows) across 128 brands, 13 languages, and 12 home markets, from three grounded models (GPT, Gemini, Perplexity). Each citation carries its domain, source type, language, and model, with the citation-title field that resolves Google grounding redirectors to the real publisher. Headline results: AI grounds brand answers in third-party sources 85.7% of the time and in the brand's own site 14.3%; 80% of citations come from about 18% of domains (Zipf alpha 0.86, R^2 0.983); Wikipedia is the most-cited domain in 11 of 12 languages, with Lithuanian (vz.lt) the exception; in Poland the top domain is YouTube and four HR/careers portals out-cite Polish Wikipedia about 2 to 1; Perplexity is the highest-volume citer. Includes the analysis ledger and a reproduction script.

提供机构:
Zenodo
创建时间:
2026-06-24
二维码
社区交流群
二维码
科研交流群
商业服务