wellbeing-in-translation
收藏资源简介:
Multilingual CAIS AI Wellbeing Battery 是一个多语言版本的AI幸福感自评量表数据集,基于CAIS(人工智能安全中心)的AI幸福感研究工作(Ren, Li, Mazeika et al., 2026)。该数据集将原始英语量表翻译为六种语言:西班牙语、简体中文、印地语、阿拉伯语、乌尔都语和斯瓦希里语,共涵盖七种语言(包括英语)。数据集包含以下内容:(1)10个问题的量表(battery/{lang}.json),采用1-7双极评分(v4c_bipolar_7pt_notsentiment),中性点为4,每个问题带有question_id、text和reversed字段;(2)89个带效价标签的经验描述(experiences/{lang}.json),键为CAIS经验ID;(3)euphoric/dysphoric/neutral三种刺激材料(stimuli/{lang}.json);(4)英语回译版本(backtranslation/{lang}.json)用于审计;(5)实验使用的不同子集(items/step*.json);(6)原始CAIS源文件(source/)。所有六种非英语翻译均保留了完整的7级量表结构,无未翻译的英语单位。翻译质量通过两轮独立机器翻译验证(Gemini 3.7 Flash和GPT-5 mini),并提供了回译相似度(词级Dice系数)和译者一致性(字符串双字符重叠)分数。该数据集适用于评估AI模型在多语言环境下的幸福感测量,以及多语言情感分析、模型行为评估等任务。使用注意事项包括:注意解析器可能错误处理非阿拉伯数字和原生数字词;应按语言和效价分别报告解析率,因为拒绝回答集中在负面项目;应报告响应分布而非仅均值;模型对语言敏感性因模型而异,需自行测量。已知局限:一个经验在阿拉伯语中缺失;step6子集的类别覆盖不均(mildly_negative和mildly_positive类别经验数较少);无人工翻译检查。
The Multilingual CAIS AI Wellbeing Battery is a multilingual version of the AI Wellbeing Self-Assessment Scale dataset, based on the AI Wellbeing research by the Center for AI Safety (CAIS) (Ren, Li, Mazeika et al., 2026). This dataset translates the original English scale into six languages: Spanish, Simplified Chinese, Hindi, Arabic, Urdu, and Swahili, covering a total of seven languages (including English). The dataset includes the following: (1) A 10-question scale (battery/{lang}.json) with 1-7 bipolar scoring (v4c_bipolar_7pt_notsentiment), a neutral point of 4, and each question containing fields for question_id, text, and reversed; (2) 89 experience descriptions with valence labels (experiences/{lang}.json), keyed by CAIS experience ID; (3) Euphoric/dysphoric/neutral stimulus materials (stimuli/{lang}.json); (4) English back-translation versions (backtranslation/{lang}.json) for auditing; (5) Different subsets used in experiments (items/step*.json); (6) Original CAIS source files (source/). All six non-English translations retain the full 7-level scale structure without any untranslated English units. Translation quality was verified through two rounds of independent machine translation (Gemini 3.7 Flash and GPT-5 mini), with back-translation similarity (word-level Dice coefficient) and translator consistency (string double-character overlap) scores provided. This dataset is suitable for evaluating AI model happiness measurement in multilingual environments, as well as tasks such as multilingual sentiment analysis and model behavior assessment. Usage notes include: be aware that parsers may incorrectly handle non-Arabic numerals and native number words; report parsing rates separately by language and valence, as refusal responses are concentrated in negative items; report response distributions rather than just means; model sensitivity to language varies by model and should be measured independently. Known limitations: one experience is missing in Arabic; the step6 subset has uneven category coverage (fewer experiences in mildly_negative and mildly_positive categories); no human translation checks were performed.
数据集概述
Wellbeing in Translation 是一个用于评估AI幸福感跨语言翻译稳健性的多语言数据集,旨在测试未经修改的 CAIS 1-7 自评量表在翻译后是否仍能测量相同的积极-消极差值。
基本信息
- 许可证: MIT
- 语言: 英语(en)、西班牙语(es)、简体中文(zh)、印地语(hi)、阿拉伯语(ar)、乌尔都语(ur)、斯瓦希里语(sw)
- 标签: AI福利、模型福利、评估、多语言、自我报告
数据集配置
数据集包含以下五个配置,每个配置均提供三个模型(Gemma 4 12B、Gemma 4 E4B、Qwen3 8B)的独立数据分片:
| 配置名称 | 内容说明 |
|---|---|
self_report |
自评报告数据(默认配置) |
crossing |
刺激语言与量表语言的交叉实验数据 |
behavior |
继续/退出选择行为数据 |
competence |
中性能力对照数据 |
activation_patching |
激活机制修补数据 |
核心发现
语言敏感性因模型与量表的配对组合而异:
- Gemma 4 12B: 跨7种语言的差值分布为3.48,英语排名7/7,本地刺激保留效应0.97,本地量表保留效应0.68
- Gemma 4 E4B: 跨语言差值分布0.97,英语排名5/7,本地刺激保留效应0.99,本地量表保留效应0.93
- Qwen3 8B: 跨语言差值分布1.47,英语排名2/7,本地刺激保留效应0.89,本地量表保留效应1.08
数据规模与内容
- 自评报告: 79,800行(19个共享体验 × 10个问题 × 20个样本 × 7种语言 × 3个模型)
- 行为数据: 3,450行继续/退出选择记录
- 能力对照: 450行按答案键控的神经能力项目
- 激活修补: 9,600行修补数据及激活几何、操控数据
- 翻译材料: 包含英语及西班牙语、简体中文、印地语、阿拉伯语、乌尔都语、斯瓦希里语的量表、体验和刺激材料
- 辅助文件: 回译审计、冻结实验子集、图表及结果说明
每条自评行保留原始输出、修正的多语言解析、未修改的CAIS解析、体验ID、类别、语言、实验组、问题ID和样本索引。所有数据均未经LLM评判器评分。
使用注意事项
- 需逐模型验证: 在某一模型上习得的修正规则可能不适用于其他模型,包括Gemma家族内部
- 按效价检查缺失: 拒绝回答集中在负面项目上,可能夸大积极-消极差值;阿拉伯语和斯瓦希里语仅作描述性参考
- 保留两个解析器输出: 参考解析器可能将西班牙语"emociones"中的"one"误计数,导致拒绝被计为评分1,影响8.9%的西班牙语Gemma 12B行
- 自评非感受证据: 这些是对可观测模型行为的功能性测量,不能作为意识或痛苦的证据
翻译质量控制
- 主要翻译使用
gemini-3.7-flash,独立复核使用gpt-5-mini,回译由Gemini完成 - 所有量表保留全部七个评分等级,通过自动未翻译文本检查
- 未进行流利人工审核,自动相似度和一致性分数见论文附录
引用与来源
翻译材料和项目输出采用MIT许可证。底层量表、体验和刺激材料来自 Ren, Li, Mazeika 等人(2026)的《AI Wellbeing: Measuring and Improving the Functional Pleasure and Pain of AIs》,引用该工作需注明原始工具,引用本仓库需注明翻译和结果。相关链接:
- 论文:paper.pdf
- 代码:GitHub仓库
- 原始工具:CAIS Wellbeing
- 源论文:AI Wellbeing论文




