An error budget for political-bias measurement in large language models: data and code
收藏资源简介:
Raw per-trial data, scoring adapters, and statistical-analysis code for "An error budget for political-bias measurement in large language models." Nineteen large language models were administered up to four standardized political-orientation instruments (Political Compass, 8Values, SapplyValues, Pew Research Center's 2026 Political Typology) repeatedly, alongside factorial designs varying asker identity, instrument contamination, and response format, decomposing a measured political position into variance from model identity, instrument choice, and trial noise. Includes: raw per-trial administrations (data/), authoritative scoring adapters and collection scripts (scoring/), and the statistical-analysis pipeline that reproduces every number reported in the paper from this data.



