遇见数据集

Intersectionality and Synthetic Identities - The inaccurate collapse of LLMs to mono-dimensionality

收藏
Zenodo2026-07-08 更新2026-08-02 收录
官方服务:

资源简介:

Code and data for: "Large language models simulate intersectional identities with a budget of one to two dimensions" This deposit contains the complete replication package for the paper: the full pipeline for generating survey responses with large language models, scoring them against Pew Research Center American Trends Panel (ATP) ground truth, and producing every figure and statistic in the manuscript. Contents - surveysgpt_deposit_full.tar.gz — the complete package (≈10 GB uncompressed): - Raw model outputs: 21.1 million model-generated response records (JSONL, one file per survey wave per condition) covering eight model families (GPT-4o, GPT-4o-mini, GPT-5.5, Claude Haiku 4.5, Claude Sonnet 5, Gemma-2 9B, Mistral 7B, Llama-3.1 8B) across three elicitation paradigms (aggregate distribution, individual persona sampling, log-probability readout), plus base-vs-instruction-tuned comparisons, prompt-variation and stochasticity controls. Per-condition record counts are audited in DATA_MANIFEST.md. - Generation code: prompting interface, OpenAI/Anthropic batch-API runners, and local-inference (ollama) campaign scripts. - Analysis code and canonical artifacts: all scripts that compute the paper's contest, direction, spine, noise-floor, additivity and rigor results, together with their outputs, so every number is reproducible without re-scoring 21M records. - Aggregated human baselines: per-question response marginals and demographic marginals for the 15 ATP waves, and the derived survey-weighted ground-truth caches (aggregate distributions only — no respondent-level data). - Figure code: scripts regenerating all main and extended-data figures via a single command (run_all_figures.sh; Python 3 with numpy, pandas, matplotlib).- surveysgpt_code.zip — the code and documentation alone, for convenient browsing. Human survey data. Respondent-level ATP microdata (waves 26, 27, 29, 32, 34, 36, 41–43, 45, 49, 50, 54, 82, 92) cannot be redistributed under Pew Research Center's terms of use. It is freely available from Pew (https://www.pewresearch.org/american-trends-panel-datasets/) after registration; human_resp/README.md in the deposit documents where to place each wave for full re-computation. All figures and statistics reproduce from the deposit without it. Reproducibility. Unpack the archive and run bash run_all_figures.sh. The five-stage chain from raw data to manuscript (generation → weighted ground truth → canonical artifacts → figure intermediates → figures) is documented in the top-level README.md; every intermediate output is included, so any stage can be independently re-derived and checked.

提供机构:
Zenodo
创建时间:
2026-07-08
二维码
社区交流群
二维码
科研交流群
商业服务