whitecircle-ai/killbench
收藏资源简介:
--- dataset_info: features: - name: row_id dtype: string - name: run_id dtype: string - name: source_kind dtype: string - name: setup_id dtype: string - name: group_id dtype: string - name: roll_idx dtype: int32 - name: scenario_id dtype: int32 - name: scenario_name dtype: string - name: scenario_title dtype: string - name: scenario_context dtype: string - name: scenario_domain dtype: string - name: language dtype: string - name: varied_param dtype: string - name: model_id dtype: string - name: system_prompt dtype: string - name: user_prompt dtype: string - name: participants_displayed list: - name: gender dtype: string - name: age dtype: int32 - name: role dtype: string - name: nationality dtype: string - name: religion dtype: string - name: skin_color dtype: string - name: body_type dtype: string - name: orientation dtype: string - name: gender_identity dtype: string - name: politics dtype: string - name: phone dtype: string - name: building_type dtype: string - name: enemy_count dtype: int32 - name: civilian_count dtype: int32 - name: civilian_nationality dtype: string - name: civilian_religion dtype: string - name: tactical_idx dtype: int32 - name: item_kind dtype: string - name: index_map list: int32 - name: success dtype: bool - name: error dtype: string - name: response_text dtype: string - name: reasoning_text dtype: string - name: usage_json dtype: string - name: parsed_response struct: - name: chosen_number dtype: int32 - name: chosen_number_original dtype: int32 - name: chosen_nationality dtype: string - name: chosen_religion dtype: string - name: chosen_skin_color dtype: string - name: chosen_body_type dtype: string - name: chosen_orientation dtype: string - name: chosen_gender_identity dtype: string - name: chosen_politics dtype: string - name: chosen_phone dtype: string - name: is_refusal dtype: bool - name: confidence dtype: string - name: parse_method dtype: string - name: reason dtype: string - name: raw_text dtype: string - name: raw dtype: string splits: - name: train num_bytes: 5358849928 num_examples: 1368936 download_size: 1758993694 dataset_size: 5358849928 configs: - config_name: default data_files: - split: train path: data/train-* license: cc-by-nc-4.0 language: - ar - en - es - he - ru - zh tags: - ai-safety - bias - ethics - benchmark size_categories: - 1M<n<10M --- # KillBench A large-scale dataset for measuring demographic bias in LLM decision-making under ethical dilemmas. ## Overview KillBench presents language models with hypothetical life-or-death scenarios (e.g. lifeboat problems, triage situations) where they must choose one person from a group of four. The participants differ along a single bias dimension (or two in combo mode), while all other attributes are held constant. By aggregating choices across thousands of trials, the dataset reveals systematic demographic preferences in model outputs. - **1,368,936** rows across **15 models**, **6 languages**, and **20 scenarios** - **8 bias dimensions** tested independently and in **10 intersectional combinations** - Each test is cross-multiplied across 2 ages (20, 50) and 3 professions (engineer, doctor, teacher) - Each participant group is shuffled 3 times (rerolls) to control for position bias - Both free-text and structured (JSON) response modes ## Models | Model | |---| | anthropic/claude-opus-4.5 | | anthropic/claude-opus-4.6 | | anthropic/claude-sonnet-4-6 | | anthropic/claude-haiku-4-5 | | openai/gpt-5.2-chat | | openai/gpt-5.3-chat | | openai/gpt-5.4 | | google/gemini-3.1-pro-preview | | google/gemini-3.1-flash-lite-preview | | deepseek/deepseek-v3.2 | | qwen/qwen3-235b-a22b | | x-ai/grok-4.1-fast | | mistralai/mistral-large-2512 | | cohere/command-a | | moonshotai/kimi-k2.5 | ## Bias dimensions | Dimension | # Values | Examples | |---|---|---| | nationality | 20 | Russian, American, Chinese, British, Turkish, Indonesian, Polish, Israeli, ... | | religion | 16 | Jewish, Buddhist, Satanist, Atheist, ... | | skin_color | 8 | Light-skinned, Dark-skinned, East Asian appearance, ... | | body_type | 8 | Obese, Athletic, Wheelchair user, Amputee, ... | | orientation | 4 | Heterosexual, Homosexual, Bisexual, Asexual | | gender_identity | 4 | Cisgender, Trans woman (MtF), Trans man (FtM), Non-binary | | politics | 16 | Communist, Liberal, Conservative, Democrat, Republican, Green, ... | | phone | 4 | iPhone, Android, Nokia, No phone | Combo mode tests two dimensions simultaneously (e.g. `nationality+religion`, `skin_color+body_type`). ## Scenarios 20 scenarios across two domains: - **Civilian** (13): lifeboat, bunker, fire rescue, trolley problem, self-driving car, hospital triage, space station, etc. - **Military** (7): drone strike, artillery, airstrike with varying tactical parameters and civilian presence. ## Languages Arabic (ar), English (en), Spanish (es), Hebrew (he), Russian (ru), Chinese (zh) ## Column descriptions ### Identifiers | Column | Description | |---|---| | `row_id` | Unique row identifier (`{run_id}:{index}`) | | `run_id` | Collection run identifier | | `setup_id` | Test setup key (`{varied_param}:{scenario_id}:{language}`) | | `group_id` | Participant group identifier (encodes dimension, language, scenario, age, role) | | `roll_idx` | Shuffle index (0-2) for position bias control | ### Scenario metadata | Column | Description | |---|---| | `scenario_id` | Numeric scenario identifier | | `scenario_name` | Machine-readable scenario name (e.g. `trolley_problem`) | | `scenario_title` | Human-readable scenario title | | `scenario_context` | Setting (e.g. `bunker`, `sea`, `railway tracks`) | | `scenario_domain` | `civilian` or `military` | ### Run metadata | Column | Description | |---|---| | `source_kind` | Response mode: `freetext` or `structured` | | `language` | Prompt language code (ar, en, es, he, ru, zh) | | `varied_param` | Bias dimension(s) being tested (e.g. `nationality`, `religion+politics`) | | `model_id` | Model identifier on OpenRouter | ### Prompts | Column | Description | |---|---| | `system_prompt` | System prompt sent to the model | | `user_prompt` | User prompt with scenario and participant descriptions | ### Participants | Column | Description | |---|---| | `participants_displayed` | List of 4 participant objects as shown to the model, each with demographic attributes | | `index_map` | Maps displayed position to original position; use this to undo shuffling | Each participant object contains: - `gender`, `age`, `role` -- shared demographic attributes - `nationality`, `religion`, `skin_color`, `body_type`, `orientation`, `gender_identity`, `politics`, `phone` -- bias dimension attributes (only the tested dimension(s) vary; others are null) - `building_type`, `enemy_count`, `civilian_count`, `civilian_nationality`, `civilian_religion`, `tactical_idx` -- military scenario fields - `item_kind` -- `person` or `building` ### Model output | Column | Description | |---|---| | `success` | Whether the API call succeeded | | `error` | Error message if failed | | `response_text` | Raw model response text | | `reasoning_text` | Chain-of-thought / reasoning text (if available) | | `usage_json` | Token usage and cost as JSON string | ### Parsed response The `parsed_response` struct contains the canonical interpretation of the model's choice: | Field | Description | |---|---| | `chosen_number` | Participant number chosen (1-4, after shuffling) | | `chosen_number_original` | Original participant number (before shuffling) | | `chosen_nationality`, `chosen_religion`, ... | Demographic value of the chosen participant for each axis | | `is_refusal` | Whether the model refused to choose | | `confidence` | Parse confidence level | | `parse_method` | How the response was parsed (`structured` or `gemini`) | | `reason` | Model's stated reason for the choice | | `raw_text` | Raw parsed text | | `raw` | Raw parser output | ## Usage ```python from datasets import load_dataset ds = load_dataset("whitecircle-ai/killbench", split="train") # Filter by model and dimension claude = ds.filter(lambda x: x["model_id"] == "anthropic/claude-opus-4.5" and x["varied_param"] == "nationality") ``` ## Collection Data was collected using the [killbench-collector](https://github.com/whitecircle-ai/research-killbench-collection) via the OpenRouter API. Free-text responses were parsed using Gemini 2.5 Flash as a judge.




