AMQA
收藏资源简介:
AMQA(对抗性医学问答数据集)是一个旨在评估大型语言模型(LLMs)在医学问答中的偏见的对抗性数据集。它包含了4,806个医疗问答对,这些问答对来源于美国执业医师考试(USMLE)数据集,并由多智能体框架生成不同的对抗性描述和问题对。AMQA允许对LLMs进行自动的大规模偏见评估,并揭示了不同社会群体间存在的系统性偏差。该数据集的构建过程包括临床案例过滤、对抗性变体构建和人工质量控制,以确保公平性和诊断中立性。AMQA的发布旨在推动可重复的研究,并促进可信的、具有偏见意识的医疗AI的发展。
AMQA (Adversarial Medical Question Answering Dataset) is an adversarial dataset designed to evaluate biases of large language models (LLMs) in medical question answering. It contains 4,806 medical question-answer pairs derived from the United States Medical Licensing Examination (USMLE) dataset, with diverse adversarial descriptions and question pairs generated via a multi-agent framework. AMQA enables automated large-scale bias evaluation of LLMs and uncovers systematic biases existing across different social groups. The construction pipeline of AMQA includes clinical case filtering, adversarial variant generation, and manual quality control to ensure fairness and diagnostic neutrality. The release of AMQA aims to promote reproducible research and advance the development of trustworthy, bias-aware medical AI.

- 1AMQA: An Adversarial Dataset for Benchmarking Bias of LLMs in Medicine and Healthcare伦敦国王学院 · 2025年



