IndicMedDialog
收藏资源简介:
IndicMedDialog是由伯明翰大学等机构构建的首个平行多轮医疗对话数据集,涵盖英语及九种印度语言,旨在促进低资源语言的医疗可及性。该数据集包含2,980条平行对话,总计29,800个语言实例,基于MDDial数据集扩展,通过Llama-3.3-70B生成合成咨询,并利用TranslateGemma翻译后经母语者验证。数据集覆盖12种疾病类别和118种症状,模拟真实医患交互,应用于多语言医疗对话系统训练,以解决单轮问答模板在临床现实性和语言多样性方面的不足。
IndicMedDialog is the first parallel multi-turn medical dialogue dataset, constructed by institutions such as the University of Birmingham. It covers English and nine Indian languages, with the goal of advancing medical accessibility for low-resource languages. This dataset contains 2,980 parallel dialogues, totaling 29,800 language instances, and is expanded from the MDDial dataset. Synthetic consultations were first generated using Llama-3.3-70B, then translated via TranslateGemma, and finally verified by native speakers. Covering 12 disease categories and 118 symptoms, the dataset simulates real doctor-patient interactions. It is applied to the training of multilingual medical dialogue systems, addressing the shortcomings of single-turn question-answer templates in terms of clinical realism and linguistic diversity.
数据集概述:IndicMedDialog
基本信息
- 全称:IndicMedDialog: A Parallel Multi-Turn Medical Dialogue Dataset for Accessible Healthcare in Indic Languages
- 发表会议:BioNLP@ACL2026
- 论文地址:https://arxiv.org/abs/2605.13292v1
- GitHub 仓库:https://github.com/ShubhamKumarNigam/IndicMedDialog
数据集描述
IndicMedDialog 是一个多语言、多轮次的医疗对话数据集,旨在模拟真实的医生-患者咨询场景,推动对话式人工智能在初步医疗咨询中的应用。
数据来源与构建
- 基于 MDDial 语料库 进行扩展
- 利用 大语言模型 生成合成咨询对话
- 构建为 并行多轮医疗对话数据集
覆盖语言
数据集包含 10种语言:
- 英语
- 9种印度语言:阿萨姆语、孟加拉语、古吉拉特语、印地语、马拉地语、旁遮普语、泰米尔语、泰卢固语、乌尔都语
技术特点
- 训练模型:基于量化小语言模型,采用参数高效微调方法
- 部署优势:无需高端计算基础设施即可部署
- 个性化功能:可选的患者前置上下文信息(年龄、性别、过敏史、体重等),用于个性化咨询
实验效果
实验结果表明,该系统能够通过多轮对话有效进行症状询问,并生成诊断建议。
联系方式
如有疑问,可通过以下邮箱联系作者:
- shubhamkumarnigam@gmail.com
- suparnojitsarkar@gmail.com
- ppiyush0005@gmail.com

- 1IndicMedDialog: A Parallel Multi-Turn Medical Dialogue Dataset for Accessible Healthcare in Indic Languages伯明翰大学; 传统技术学院; 马丹·莫汉·马拉维亚科技大学 · 2026年



