Cross-Language AI Brand Reputation: A 66-Brand, 12-Language Dataset of LLM-Constructed Reputation (2026)
收藏资源简介:
Data and analysis code behind the study "The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages." Three grounded large language models (GPT-5.4, Google Gemini 3.1 Pro, Perplexity Sonar Pro) were asked about 66 brands from eleven Northern, Baltic, and Central European home markets, in twelve languages spanning four families (Germanic, Uralic, Baltic, Slavic), producing 35,640 grounded responses, with a 20-brand Central and Eastern European companion cohort adding a five-iteration dice-roll stability protocol. The dataset includes per-response sentiment, 196,020 cross-language BGE-M3 embedding similarities, 150,093 source attributions, per-brand-by-language measures, the stability cells, and a reproduction script. Headline results: mean cross-language cosine similarity 0.825 (same family 0.844 vs cross-family 0.820); sentiment varies by language (ANOVA F=268.5, eta^2=0.077); language clustering recovers the Slavic and Baltic families (cophenetic 0.915); moving to a brand's home language raises recommendation share by 0.80 for local champions versus 0.15 for global multinationals (t=-8.84, p<0.001); response stability is governed more by model than language (eta^2 0.319 vs 0.011).



