MedArabiQ
收藏资源简介:
MedArabiQ是一个包含七个阿拉伯语医疗任务的基准数据集,涵盖了多个专业领域,包括选择题、填空题和医患问答。该数据集通过过去的医疗考试和公开可用的数据集构建而成,旨在评估大型语言模型在阿拉伯语医疗领域的性能。数据集包含700条数据,通过多种方式修改以评估不同LLM的能力,包括偏见缓解。该数据集的发布为未来研究提供了基础,旨在评估和增强LLM的多语言能力,以确保在医疗保健中使用生成式AI的公平性。
MedArabiQ is a benchmark dataset encompassing seven Arabic medical tasks across multiple specialized medical domains, including multiple-choice questions, fill-in-the-blank questions, and doctor-patient Q&A. Constructed from past medical examinations and publicly available datasets, this dataset aims to evaluate the performance of large language models (LLMs) in the Arabic medical domain. It contains 700 data instances, with modifications implemented via multiple approaches to assess the capabilities of various LLMs, including bias mitigation. The release of MedArabiQ provides a foundational resource for future research focused on evaluating and enhancing the multilingual capabilities of LLMs, with the ultimate goal of ensuring fairness in the application of generative AI in healthcare.
MedArabiQ: 阿拉伯语医疗任务大型语言模型基准测试数据集
概述
- 目的:评估大型语言模型(LLMs)在阿拉伯语医疗领域的表现
- 特点:
- 包含7种阿拉伯语医疗任务
- 涵盖多种专业领域和问题格式
- 基于医学考试和公开资源构建
- 特别关注偏见缓解评估
任务类型
- 多项选择题 - 医学知识评估
- 多项选择题 - 医疗环境中的偏见评估
- 填空(提供选项)
- 填空(不提供选项)
- 医患问答(QA)
- 带语法错误纠正的QA
- 经LLM修改的QA
技术细节
- 评估模型:包含GPT-4o、Claude 3.5-Sonnet和Gemini 1.5等8种先进LLM
- 数据格式:CSV
- 内容组成:任务描述、输入提示和标准答案
应用价值
- 为多语言医疗AI模型评估提供基准
- 促进医疗AI公平性和可扩展性发展
- 支持未来多语言医疗AI研究

- 1MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks纽约大学阿布扎比分校, 阿拉伯联合酋长国 · 2025年



