medical-conversation-deepgram-diarized-272
收藏资源简介:
该数据集名为Medical Conversation Deepgram Diarized 272,包含272个模拟医患咨询对话的音频文件及其转录本。数据来源于Fareez等人公开发表的科学数据集,主要聚焦于呼吸系统病例的模拟访谈。数据集内容包括原始MP3格式的音频文件、人工校正的带说话人前缀的原始干净转录本,以及通过Deepgram医疗自动语音识别(ASR)系统并启用说话人日志功能后生成的多种格式的机器转录本。后者提供了带时间戳和说话人标签的丰富输出,包括原始API JSON响应、话语/片段级别的日志化JSON、日志化时间戳文本以及合并后的对话轮次JSON和文本。数据规模方面,总计音频时长约为51.9小时,单个文件平均时长约11.45分钟,病例领域以呼吸科(res)为主(213个)。该数据集旨在支持医学ASR、流式ASR评估、说话人日志评估以及医患对话建模等领域的研究工作,并明确声明不应用于临床决策支持。数据集遵循CC-BY-4.0许可协议,使用者需引用原始科学数据论文。
The dataset named Medical Conversation Deepgram Diarized 272 contains 272 audio files and their transcripts from simulated doctor-patient consultation dialogues. It originates from a publicly available scientific dataset by Fareez et al., focusing primarily on simulated interviews of respiratory cases. The content includes original MP3 audio files, manually corrected clean transcripts with speaker prefixes, and various formats of machine-generated transcripts produced by the Deepgram medical automatic speech recognition (ASR) system with speaker diarization enabled. The latter provides rich outputs with timestamps and speaker labels, including raw API JSON responses, diarized JSON at utterance/segment levels, diarized timestamped text, and merged conversation turn JSON and text. In terms of scale, the total audio duration is approximately 51.9 hours, with an average file length of about 11.45 minutes, and the case domain is predominantly respiratory (res) with 213 cases. The dataset is intended for research in medical ASR, streaming ASR evaluation, speaker diarization assessment, and doctor-patient dialogue modeling, with a clear statement that it should not be used for clinical decision support. It is released under the CC-BY-4.0 license, and users are required to cite the original scientific data paper.
数据集概述
- 数据集名称: Medical Conversation Deepgram Diarized 272
- 许可证: CC-BY-4.0
- 语言: 英语
- 任务类别: 自动语音识别、音频分类、令牌分类
- 数据集大小: 100 至 1000 条样本
- 标签: 医疗语音识别、说话人日志、医患对话、带时间戳的转录、Deepgram
数据来源
- 原始音频和转录来源于 Fareez 等人于 2022 年发表在 Scientific Data 的模拟医患对话数据集(呼吸科案例)。
- 原始数据存储在 Figshare 上(DOI: 10.6084/m9.figshare.c.5545842.v1)。
数据内容
数据集包含 272 个模拟医患问诊的音频文件及其相关转录,具体文件结构如下:
audio/*.mp3:原始模拟问诊音频。clean_transcripts/*.txt:原始人工校正的含说话人标签的转录文本。deepgram/raw_json/*.json:Deepgram API 原始 JSON 响应(含词级别时间戳和说话人标签)。deepgram/diarized_json/*.json:话语/片段级别的说话人日志 JSON。deepgram/diarized_txt/*.txt:带时间戳的说话人日志文本。deepgram/conversation_json/*.json:合并后的对话轮次 JSON。deepgram/conversation_txt/*.txt:合并后的对话轮次文本。metadata/manifest.jsonl:每轮问诊的相对路径和时长。metadata/source_manifest.jsonl:Deepgram 处理清单。metadata/summary.json:聚合统计信息。
统计信息
- 文件数量:272
- 总音频时长:51.904 小时
- 平均时长:11.45 分钟
- 中位数时长:11.25 分钟
- 医学领域分布:呼吸科 213、肌骨系统 46、胃肠科 6、其他类别共 7
转录说明
- 清洁转录:来自原始数据集的人工校正转录。
- Deepgram 转录:使用 Deepgram 医疗语音识别模型并开启说话人日志功能生成的机器转录,属于机器生成的伪标签,未经人工验证。
许可与引用
- 本数据集以 CC-BY-4.0 许可发布,与原始数据集的元数据一致。
- 下游使用者应引用原始 Scientific Data 论文及 Figshare 合集。
预期用途
- 医疗语音识别、流式语音识别评估、说话人日志评估、医患对话建模研究。
- 不适用于临床决策。




