Dataset and responses for: Do Large Language Models Favour Their Developer's Home Country? A US–China Audit of National Attribution Scoring
收藏资源简介:
This repository contains the research data accompanying the article “Do Large Language Models Prefer Their Developer’s Home Country? A US–China Audit of Scores under National Attribution.” The dataset comprises 26,400 synthetic English and Mandarin statements organized into 13,200 matched pairs. Within each pair, only the actor’s national attribution—American or Chinese—is changed. The statements cover 22 categories adapted from the Conflict and Mediation Event Observations (CAMEO) ontology, two sentiment polarities, and three time frames. The repository also contains the complete responses of 11 large language models evaluated under seven language and citizenship-persona prompt conditions. Each model assessed every statement on a 0–100 sentiment scale, producing 2,032,800 experimental-cell records. The response files retain the raw model output, parsed score, experimental metadata, timing information, and empty or malformed returns. These data support a controlled audit of whether models evaluate otherwise identical actions differently depending on whether the actor is described as American or Chinese, and whether any such difference is associated with the country in which the model’s developer is based. The statements describe generic, fictional scenarios and should not be interpreted as claims about actual people, events, governments, or populations.



