ArogyaBodha
收藏资源简介:
ArogyaBodha是由印度理工学院等机构构建的大规模多语言多模态医疗推理数据集,涵盖英语及七种主要印度语言。该数据集包含40,857个样本,整合了八个异构医疗数据源,覆盖31个身体系统和21个临床领域,每个样本均包含专家验证的医疗查询与明确答案。数据集通过GPT-4o-mini和Gemini-2.5-Flash进行少样本提示生成与严格过滤,并采用Gemini-2.5-Pro进行高质量翻译以确保临床语义完整性。该数据集旨在解决低资源印度语言环境下多模态医疗推理的评估瓶颈,推动医疗AI在多元化语言场景中的公平应用。
ArogyaBodha is a large-scale multilingual multimodal medical reasoning dataset developed by institutions including the Indian Institute of Technology and other relevant organizations, supporting English and seven major Indian languages. The dataset comprises 40,857 samples compiled from eight heterogeneous medical data sources, spanning 31 body systems and 21 clinical domains, with each sample containing expert-validated medical queries and definitive, unambiguous answers. The dataset was generated via few-shot prompting and rigorous filtering using GPT-4o-mini and Gemini-2.5-Flash, and underwent high-quality translation with Gemini-2.5-Pro to ensure the integrity of clinical semantics. This dataset aims to address the evaluation bottleneck of multimodal medical reasoning in low-resource Indian language contexts, and promote the equitable application of medical AI across diverse linguistic scenarios.
数据集概述:ArogyaBodha
ArogyaBodha 是一个大规模多语言多模态医学问答数据集,由 ArogyaSutra 项目发布,旨在推动印度语言环境下的多模态医学推理。
核心规模
- 总样本数:40,857 个经过专家验证的问答样本
- 数据来源:8 个异构医学来源
- 覆盖范围:31 个身体系统、6 种成像模态、21 个临床领域
- 语言:英语及 7 种主要印度语言(阿萨姆语、孟加拉语、印地语、马拉地语、旁遮普语、泰米尔语、泰卢固语)
质量控制
- 所有样本经过专家验证(Expert Verification)
- 使用 COMET-QA 进行质量控制
发布信息
- 该数据集与多智能体框架 ArogyaSutra 及微调模型 arogyasutraV1E3.5(基于 Qwen2.5-VL-7B)一同发布
- 论文已被 IJCAI 2026 接收
- 主页:https://iitp-cse.github.io/ArogyaSutra/
- 模型与数据集获取:🤗 Hugging Face 链接(来源于页面内容中的“🤗 Models & Datasets”按钮,实际地址需参考页面)

- 1ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages印度理工学院巴特那分校; 印度理工学院坎普尔分校; 普拉萨纳德布女子学院 · 2026年



