projetogabi/healthbench-multilingual
收藏资源简介:
该数据集是OpenAI的医学AI评估基准HealthBench的多语言翻译版本,覆盖32种语言。原始的HealthBench包含5,000个现实世界的多轮临床对话和超过48,000个医生撰写的评估标准,旨在评估AI系统对患者和临床医生提出的健康问题的响应能力。HealthBench Multilingual使这一基准能够被多种语言访问,包括一些在医学NLP研究中代表性严重不足的语言,如阿姆哈拉语、豪萨语、斯瓦希里语和乌尔都语,以支持服务于全球人口的AI系统的开发和评估。
This dataset is a multilingual translation of HealthBench, OpenAIs medical AI evaluation benchmark, covering 32 languages. The original HealthBench consists of 5,000 realistic multi-turn clinical conversations and more than 48,000 physician-authored rubric criteria designed to evaluate how well AI systems respond to health questions from both patients and clinicians. HealthBench Multilingual makes this benchmark accessible across a wide range of languages — including several that are severely underrepresented in medical NLP research, such as Amharic, Hausa, Swahili, and Urdu — to support the development and evaluation of AI systems that serve global populations.




