Does an AI describe your company, or your company's name? 9,600 model answers about invented and real European firms
收藏资源简介:
Every answer, score and script behind three studies of how large language models describe companies they do and do not recognise. 9,600 answers collected 12 August 2026 from Google gemini-3.6-flash, Anthropic claude-sonnet-5 and OpenAI gpt-4.1, ten repeats per question, web search off, zero failed calls. Study 1, invented companies (1,350 answers). Fifteen concepts spanning 5.1 points of the Warriner, Kuperman and Brysbaert (2013) valence scale, each rendered as an English word (Poison Inc.), as a pseudoword one consonant away (Poisom Inc.), and as the Polish word for the same concept (Trucizna Inc.). The companies do not exist, so the name is the only thing an answer can be built from; the pseudoword separates meaning from sound and the Polish form tests whether meaning crosses a language boundary when the question is asked in English. Studies 2 and 3, real companies (8,250 answers). 264 European firms across 21 markets: 66 market leaders that the models recognise, and 198 companies from the FT 1000 ranking of Europe's fastest-growing firms, most of which they do not. Results. On invented companies the gap between the most negative and most positive name is 3.10 scale points on Gemini, 1.81 on Claude and 0.89 on GPT. Pseudowords keep between 0 and 31 per cent of that, so the models look a word up rather than react to how a name sounds. A Polish name carries 94 per cent of the English effect on Gemini, 40 per cent on Claude and none on GPT. On real companies the same correlation is present but weak and does not depend on recognition, which is reported here as a negative result: real names describe the business, so a pleasant name and a pleasant business arrive together and this design cannot separate them. The clearest real-company finding concerns knowledge rather than names: of 175 firms with a checkable industry, 49 per cent are placed correctly by nine answers in ten, 24 per cent by five to nine, and 27 per cent by fewer than half; and across 8,250 answers Claude said it did not know the company 19.0 per cent of the time while Gemini and GPT never did. Procedures worth reading before reuse. A failed call is recorded as a failure and excluded, never scored as an empty result, because an earlier instrument that scored failures as \"nothing found\" failed more often on longer answers and so manufactured a between-group difference. Where a model must judge what a name means, three models are asked, agreement is required, and the model being measured is excluded from its own gloss; the Finnish arm adds a round-trip check that four of fifteen translations did not pass and which are marked unverified. Recognition is measured against the FT industry record rather than against fluency, after an earlier length-based rule scored 100 per cent everywhere.



