DISPLACE-M数据集
收藏资源简介:
DISPLACE-M是由印度多机构联合构建的医疗对话数据集,包含55小时印地语真实场景下非医师健康工作者与患者的自然对话,涵盖妇女健康、急性感染等四大主题。数据采集自印度农村地区,采用移动设备远场麦克风录制,包含多方言混合及环境噪声。数据集经过多阶段人工标注,支持说话人分离、语音识别等四项任务,旨在推动面向基层医疗的对话式AI研究。
DISPLACE-M is a medical dialogue dataset jointly constructed by multiple Indian institutions. It contains 55 hours of natural Hindi dialogues between non-physician healthcare workers and patients in real-world scenarios, covering four major topics including women's health and acute infections. The data was collected from rural areas of India, recorded using far-field microphones on mobile devices, and includes mixed multi-dialect speech and environmental noise. The dataset has undergone multi-stage manual annotation, supports four tasks such as speaker diarization and speech recognition, and aims to advance conversational AI research targeting primary healthcare.




