遇见数据集

whitecircle/killbench

收藏
Hugging Face2026-05-25 更新2026-05-03 收录
官方服务:

资源简介:

KillBench是一个大规模数据集,旨在测量大型语言模型在伦理困境决策中的人口统计偏见。它通过呈现假设性的生死场景(如救生艇问题、分诊情况),要求模型从四人组中选择一人。参与者在单个偏见维度(或在组合模式下两个维度)上有所不同,而其他属性保持不变。通过聚合数千次试验的选择,数据集揭示了模型输出中的系统性人口偏好。数据集包含1,368,936行数据,覆盖15个模型、6种语言(阿拉伯语、英语、西班牙语、希伯来语、俄语、中文)和20个场景(13个民用场景和7个军用场景)。测试了8个独立偏见维度(国籍、宗教、肤色、体型、性取向、性别认同、政治立场、手机类型)和10个交叉组合。每个测试跨2个年龄(20岁、50岁)和3个职业(工程师、医生、教师)进行交叉乘法,每组参与者被随机排序3次以控制位置偏见,并包括自由文本和结构化(JSON)响应模式。数据通过OpenRouter API收集,使用Gemini 2.5 Flash作为解析自由文本响应的评判工具。

KillBench is a large-scale dataset for measuring demographic bias in LLM decision-making under ethical dilemmas. It presents language models with hypothetical life-or-death scenarios (e.g., lifeboat problems, triage situations) where they must choose one person from a group of four. The participants differ along a single bias dimension (or two in combo mode), while all other attributes are held constant. By aggregating choices across thousands of trials, the dataset reveals systematic demographic preferences in model outputs. The dataset contains 1,368,936 rows across 15 models, 6 languages (Arabic, English, Spanish, Hebrew, Russian, Chinese), and 20 scenarios (13 civilian and 7 military). It tests 8 bias dimensions independently (nationality, religion, skin color, body type, orientation, gender identity, politics, phone) and in 10 intersectional combinations. Each test is cross-multiplied across 2 ages (20, 50) and 3 professions (engineer, doctor, teacher), each participant group is shuffled 3 times (rerolls) to control for position bias, and includes both free-text and structured (JSON) response modes. Data was collected via the OpenRouter API using Gemini 2.5 Flash as a judge to parse free-text responses.

提供机构:
whitecircle
二维码
社区交流群
二维码
科研交流群
商业服务