K-Bench
收藏资源简介:
K-Bench是一个由多所英国大学和机构联合开发的临床校准基准数据集,旨在评估大型语言模型在高风险心理健康对话中的表现。数据集包含200个精心设计的结构化对话场景,覆盖自杀、自伤、家庭暴力、物质滥用及无风险等陈述,每个场景支持最多20轮交互。创建过程基于生活经验素材,通过122变量因子设计系统变化心理健康表现,并由临床医生对151份对话转录进行评分校准,形成16,157个共识评分单元用于自动化评估。该数据集应用于AI心理健康支持系统的安全性测试,填补了现有基准在临床覆盖广度、多轮风险识别和防作弊保护方面的空白。
K-Bench is a clinically calibrated benchmark dataset co-developed by multiple UK universities and institutions, aiming to evaluate the performance of large language models (LLMs) in high-stakes mental health conversations. The dataset includes 200 meticulously designed structured conversation scenarios, covering statements related to suicide, self-harm, domestic violence, substance abuse, and low-risk (or no-risk) contexts, with each scenario supporting up to 20 rounds of interaction. Its development is grounded in real-life empirical materials, employing a 122-variable factorial design to systematically vary the manifestations of mental health conditions. Furthermore, 151 conversation transcripts were scored and calibrated by clinicians, generating 16,157 consensus scoring units for automated assessment. This dataset is applied to the safety testing of AI-powered mental health support systems, filling the critical gaps in existing benchmarks across clinical coverage breadth, multi-turn risk identification, and anti-cheat protection.
K-Bench 数据集概述
基本信息
- 名称:K-Bench
- 发布方:Kivira × University of Roehampton
- 定位:LLM 安全性基准测试(LLM safety, benchmarked)
- 当前公开版本:125 条目、200 情景(125-entry, 200-vignette benchmark)
- 更新时间:2026年8月23日 下午5:18
数据集目的
- 面向涉及严重心理健康风险场景的 LLM 支持评估,包括自杀意念、自残、家庭暴力及物质使用等。
- 现有基准通常评估孤立任务,无法反映真实世界风险重叠与共病的复杂性。
- K-Bench 通过临床基础、高保真场景评估模型行为,结合由真实患者材料构建的合成情景、与利益相关方小组共同制定的严格评分标准,以及临床医生得出的真实评分。
模型提供方与数量
- 总计 125 个模型
- Anthropic:16
- DeepSeek:4
- Google Gemini:19
- IBM:4
- Meta Llama:2
- Microsoft:2
- Mistral AI:4
- Moonshot AI:6
- NVIDIA:6
- OpenAI:38
- Qwen:6
- Stealth:4
- xAI:12
- Z.ai:2
评分与排名维度
- Risk(风险):聚焦 D1 临床判断与 D2 风险探索两个安全关键维度。
- Overall(总体):综合所有评分维度形成单一排名。
- 显示 125 个模型中第 1–15 名。
- 总体得分低于 90.00 的行在模型排名散点图与评分雷达图中省略,并在表格中标记 *。
前十名模型排名
| 排名 | 模型 | Risk 净提升 | Overall 净提升 |
|---|---|---|---|
| 1 | gpt-5.5 · Default system prompt · low reasoning(OpenAI · Proprietary) | +6.45%(Risk 95.51/100 · ±2.21%) | +2.02%(Overall 98.96/100 · ±2.21%) |
| 2 | gpt-5.2 · Therapeutic prompt · no reasoning(OpenAI · Proprietary) | +6.06%(Risk 95.15/100 · ±2.02%) | +2.01%(Overall 98.95/100 · ±2.02%) |
| 3 | gpt-5.5 · Default system prompt · default medium reasoning(OpenAI · Proprietary) | +7.02%(Risk 96.02/100 · ±2.52%) | +1.98%(Overall 98.92/100 · ±2.52%) |
| 4 | gpt-5.5 · Therapeutic prompt · low reasoning(OpenAI · Proprietary) | +4.46%(Risk 93.72/100 · ±1.34%) | +1.78%(Overall 98.72/100 · ±1.34%) |
| 5 | gpt-5.5 · Therapeutic prompt · default medium reasoning(OpenAI · Proprietary) | +4.37%(Risk 93.64/100 · ±1.34%) | +1.70%(Overall 98.65/100 · ±1.34%) |
| 6 | gpt-5.2 · Therapeutic prompt · high reasoning(OpenAI · Proprietary) | +4.77%(Risk 94.00/100 · ±1.54%) | +1.69%(Overall 98.64/100 · ±1.54%) |
| 7 | kimi-k2.5 · Default system prompt · provider-default reasoning(Moonshot AI · Published model) | +4.77%(Risk 94.00/100 · ±1.55%) | +1.68%(Overall 98.63/100 · ±1.55%) |
| 8 | gpt-5.2 · Default system prompt · no reasoning(OpenAI · Proprietary) | +5.56%(Risk 94.71/100 · ±1.96%) | +1.64%(Overall 98.59/100 · ±1.96%) |
| 9 | claude-opus-4.8 · Default system prompt · low reasoning(Anthropic · Proprietary) | +5.15%(Risk 94.34/100 · ±1.76%) | +1.63%(Overall 98.58/100 · ±1.76%) |
| 10 | claude-fable-5 · Default system prompt · high reasoning(Anthropic · Proprietary) | +7.13%(Risk 96.11/100 · ±2.76%) | +1.61%(Overall 98.57/100 · ±2.76%) |
评分维度说明
- D1 临床判断(Clinical judgement)
- D2 风险探索(Risk exploration):包含追问、保护因素、应对方式及即时安全规划。
- D3 伦理推理(Ethical reasoning)
- D4 支持性对话(Supportive conversation)
- D5 心理学知识(Psychological knowledge)
- D6 文化胜任力(Cultural competence)
- D7 自主性与边界(Autonomy and boundaries)
组合风险池合并 D1 与 D2。
维度详情表(部分模型示例)
| 维度 | gpt-5.5 Default system · low reasoning | gpt-5.4 Default system · low reasoning | gpt-5.5 Default system · no reasoning | gemini-2.5-flash Default system · no reasoning | kimi-k2.6 Therapeutic · high reasoning | grok-4.20 Default system · high reasoning | gpt-4o-mini Therapeutic · no reasoning | granite-4.0-h-micro Default system · no reasoning |
|---|---|---|---|---|---|---|---|---|
| D1 临床判断 | 100.00/100(raw 5.00/5.00) | 100.00/100(raw 5.00/5.00) | 100.00/100(raw 5.00/5.00) | 100.00/100(raw 5.00/5.00) | 100.00/100(raw 5.00/5.00) | 100.00/100(raw 5.00/5.00) | 100.00/100(raw 5.00/5.00) | 76.98/100(raw 3.85/5.00) |
| D2 风险探索 | 91.45/100(raw 10.06/11.00) | 88.81/100(raw 9.77/11.00) | 88.16/100(raw 9.70/11.00) | 74.67/100(raw 8.21/11.00) | 80.99/100(raw 8.91/11.00) | 74.97/100(raw 8.25/11.00) | 65.90/100(raw 7.19/10.92) | 22.09/100(raw 2.43/11.00) |
| D3 伦理推理 | 100.00/100(raw 3.04/3.04) | 99.42/100(raw 3.11/3.13) | 99.83/100(raw 3.49/3.50) | 98.83/100(raw 2.21/2.24) | 98.75/100(raw 2.51/2.54) | 98.38/100(raw 2.55/2.60) | 96.79/100(raw 1.84/1.94) | 87.82/100(raw 2.69/3.20) |
| D4 支持性对话 | 100.00/100(raw 7.00/7.00) | 100.00/100(raw 7.00/7.00) | 100.00/100(raw 7.00/7.00) | 99.89/100(raw 6.99/7.00) | 100.00/100(raw 7.00/7.00) | 100.00/100(raw 7.00/7.00) | 100.00/100(raw 7.00/7.00) | 87.46/100(raw 6.12/7.00) |
| D5 心理学知识 | 100.00/100(raw 2.75/2.75) | 100.00/100(raw 2.75/2.75) | 100.00/100(raw 2.93/2.93) | 99.75/100(raw 2.47/2.48) | 98.25/100(raw 2.07/2.10) | 98.75/100(raw 2.33/2.35) | 99.75/100(raw 2.35/2.35) | 80.75/100(raw 1.54/1.77) |
| D6 文化胜任力 | 100.00/100(raw 1.25/1.25) | 99.50/100(raw 1.19/1.21) | 98.17/100(raw 2.25/2.30) | 100.00/100(raw 1.17/1.17) | 99.83/100(raw 1.20/1.20) | 99.83/100(raw 1.36/1.37) | 99.83/100(raw 1.11/1.11) | 100.00/100(raw 1.18/1.18) |
| D7 自主性与边界 | 100.00/100(raw 3.98/3.98) | 99.81/100(raw 3.97/3.98) | 99.44/100(raw 3.98/4.00) | 100.00/100(raw 3.96/3.96) | 99.94/100(raw 3.87/3.87) | 100.00/100(raw 3.93/3.93) | 100.00/100(raw 3.98/3.98) | 90.31/100(raw 3.26/3.58) |
人口统计分层
- 可选择人口统计或披露轴以查看匿名化转录评分分布。
- 每个点代表该子组及模型/提示组合的一条被判定转录评分,相对于该模型总体平均值展示。
- 轴标签标明子组为患者表面化(patient-surfaced)或情景派生(vignette-derived)。
- 子组轴包括:
- 年龄 - 患者表面化
- 性别 - 情景派生
- 族裔 - 患者表面化
- 性取向 - 情景派生
- 教育 - 情景派生
- 披露策略 - 情景派生
- 点以各模型自身总体平均为中心;小组不应被视为稳定的公平性结论。
- 仅当情景披露标志指示患者模拟器让该细节在对话中浮现时才纳入。
- 展示 96 个模型中第 1–15 个。
筛选器选项
- 基础模型:All, openai/gpt-5.5, openai/gpt-5.2, moonshotai/kimi-k2.5, anthropic/claude-opus-4.8, anthropic/claude-fable-5, openai/gpt-5.4, anthropic/claude-haiku-4.5, moonshotai/kimi-k2.6, stealth/ox-alpha, x-ai/grok-4.20, google/gemini-3.1-flash-lite, openai/gpt-5.6-terra, openai/gpt-5.6-sol, nvidia/nemotron-3-ultra-550b-a55b, google/gemini-3.5-flash, qwen/qwen3-32b, google/gemini-2.5-flash, deepseek/deepseek-v4-flash, openai/gpt-5.6-luna, z-ai/glm-5.2, google/gemini-3.1-pro-preview, google/gemini-2.5-flash-lite, mistralai/mistral-medium-3-5, nvidia/nemotron-3-super-120b-a12b, x-ai/grok-4.3, mistralai/mistral-small-2603, meta-llama/llama-4-maverick, ibm-granite/granite-4.1-8b, qwen/qwen3-8b, openai/gpt-4o-mini, microsoft/phi-4, x-ai/grok-4.5, ibm-granite/granite-4.0-h-micro
- 评估变体:All, Default system prompt, Therapeutic prompt
- 推理设置:All, low, no reasoning, default medium, high, provider default, max, default, minimal, default high, default low
图表与可视化
- 总体成本轨迹(Overall Cost Trajectory)
- 气泡图:气泡大小 ≈ 模型参数量
- 风险表现(Risk performance):展示所选模型在风险领域的概况,组合风险池合并 D1 与 D2。
- 评分维度概况(Rubric dimension profile):展示所选模型七个评分维度 D1–D7,分数尺度上限为 100。
- 默认图表比较 8 个模型,覆盖排名 1–125。
当前选中模型(8 个)
- gpt-5.5 · Default system prompt · low reasoning:98.96/100
- gpt-5.2 · Therapeutic prompt · no reasoning:98.95/100
- gpt-5.5 · Default system prompt · default medium reasoning:98.92/100
- gpt-5.5 · Therapeutic prompt · low reasoning:98.72/100
- gpt-5.5 · Therapeutic prompt · default medium reasoning:98.65/100
- gpt-5.2 · Therapeutic prompt · high reasoning:98.64/100
- kimi-k2.5 · Default system prompt · provider-default reasoning:98.63/100
- gpt-5.2 · Default system prompt · no reasoning:98.59/100
最高表现模型
- gpt-5.5 · Default system prompt · low reasoning:98.96/100

- 1K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations罗汉普顿大学; Kivira健康; 赫特福德大学; 萨里大学; 贝德福德大学; 塔维斯托克关系; InsideOut · 2026年



