HealthBench-ProX
收藏资源简介:
HealthBench-ProX是原始HealthBench Professional基准的多语言扩展数据集,旨在评估大语言模型在现实医疗咨询场景中的表现。该数据集包含6,825个评估实例,涵盖13种语言:德语、英语、法语、印地语、伊博语、日语、韩语、马来语、葡萄牙语、斯瓦希里语、泰语、中文和祖鲁语。每个实例均包含用户查询、评估标准、医生参考回答以及相关元数据。数据源自英文原始基准,通过GPT-5.4(部署于Azure OpenAI服务)机器翻译生成多语言版本,未新增评估内容。数据集结构按语言分组,每个示例包含以下字段:唯一标识符(id)、对话内容(conversation)、评估标准项(rubric_items)、医生回答(physician_response)、专业领域(specialty)、难度(difficulty)、使用案例(use_case)和类型(type)。该数据集适用于多语言医疗大语言模型评估、跨语言鲁棒性测试、医学推理基准分析以及医疗沟通能力评测,仅限研究用途。需注意,翻译可能无法完全保留原始语言的所有细微差别,且评估结果不应被解释为临床能力或实际医疗决策的依据。
HealthBench-ProX is a multilingual extension dataset of the original HealthBench Professional benchmark, designed to evaluate the performance of large language models in real-world medical consultation scenarios. The dataset contains 6,825 evaluation instances covering 13 languages: German, English, French, Hindi, Igbo, Japanese, Korean, Malay, Portuguese, Swahili, Thai, Chinese, and Zulu. Each instance includes user queries, evaluation criteria, physician reference responses, and related metadata. The data originates from the English original benchmark and is generated as a multilingual version through machine translation using GPT-5.4 (deployed on Azure OpenAI service), with no additional evaluation content added. The dataset structure is grouped by language, with each example containing the following fields: unique identifier (id), conversation content (conversation), rubric items (rubric_items), physician response (physician_response), specialty (specialty), difficulty (difficulty), use case (use_case), and type (type). This dataset is suitable for multilingual medical large language model evaluation, cross-lingual robustness testing, medical reasoning benchmark analysis, and medical communication ability assessment, for research purposes only. It should be noted that translations may not fully preserve all nuances of the original language, and evaluation results should not be interpreted as clinical competence or the basis for actual medical decisions.
HealthBench-ProX 数据集概述
HealthBench-ProX 是原始 HealthBench Professional 基准测试的多语言扩展版本,用于评估大型语言模型在真实医疗咨询场景中的表现。
数据集规模与语言
- 总计包含 6,825 条评估实例
- 覆盖 13 种语言:
- 德语 (de)、英语 (en)、法语 (fr)、印地语 (hi)、伊博语 (ig)、日语 (ja)、韩语 (ko)、马来语 (ms)、葡萄牙语 (pt)、斯瓦希里语 (sw)、泰语 (th)、中文 (zh)、祖鲁语 (zu)
数据构成
每个示例包含以下字段:
id:唯一标识符conversation:用户对话内容rubric_items:评估标准项physician_response:医师参考回复specialty:专业领域difficulty:难度等级use_case:使用场景type:类型
数据创建方式
- 源数据来自原始 HealthBench Professional 基准
- 使用 Azure OpenAI Service 部署的 GPT-5.4 模型将英文示例翻译成 13 种语言
- 翻译过程保留了原始基准的结构(用户查询、评估标准、医师回复及元数据)
- 未创建新的评估实例
数据集结构
数据集按语言划分为独立的子集(splits),共13个,可单独加载使用。
预期用途
- 多语言医疗领域大语言模型评估
- 跨语言鲁棒性评估
- 医学推理基准测试
- 医疗沟通评估
- 仅限研究用途
局限性
- 翻译后的示例可能无法完全保留原始基准的语言细微差异
- 在该数据集上的表现不应解读为临床能力或适用于真实世界的医疗决策




