LLM-multitudes-neurips-2026/LLM-Multitudes
收藏资源简介:
该数据集名为LLMs Contain Multitudes,包含来自五个大型语言模型(Llama-8B-Instruct、Llama-70B-Instruct、Mistral Small 4、Qwen-3-30B-MoE、Claude Sonnet 4.6)在五种不同部署上下文(neutral、news、reddit、school、vlog)下的成对偏好和效用判断。数据用于支持NeurIPS 2026匿名提交的论文,研究部署上下文如何重塑模型级偏好和价值观。数据集包括两个主要实验:国家实验(涉及15个国家、6个特征、105对国家组合)和效用实验(涉及50个结果、1225对组合),并包含消融研究(如替代提示、无推理强制选择、温度扫描)。总共有1.1百万行数据,约10亿生成令牌的模型推理,以Parquet格式(42个文件,1.23 GB)提供,采用CC BY 4.0许可证。数据集旨在用于AI安全/对齐研究,特别是上下文依赖性的分析,并可用于复制和扩展论文实验。
This dataset is named LLMs Contain Multitudes. It contains pairwise preference and utility judgments from five large language models (LLMs: Llama-8B-Instruct, Llama-70B-Instruct, Mistral Small 4, Qwen-3-30B-MoE, Claude Sonnet 4.6) across five distinct deployment contexts (neutral, news, reddit, school, vlog). This dataset supports the anonymous NeurIPS 2026 submission paper investigating how deployment contexts reshape model-level preferences and values. The dataset encompasses two core experiments: the National Experiment (covering 15 countries, 6 features, and 105 country pair combinations) and the Utility Experiment (covering 50 outcomes and 1225 pair combinations), along with ablation studies such as alternative prompts, forced-choice without reasoning, and temperature scanning. In total, it contains 1.1 million rows of data, corresponding to approximately 1 billion model inference-generated tokens. The dataset is provided in Parquet format (42 files, 1.23 GB) under the CC BY 4.0 license. This dataset is intended for AI safety/alignment research, particularly for analyses of contextual dependency, and can be used to replicate and extend the paper’s experiments.



