MedTranscripts - A multimodal dataset of Spanish medical videos and time-aligned transcripts
收藏资源简介:
MedTranscripts is a dataset of 30 hours of medical videos and audios in Spanish, time-aligned with the corresponding transcripts. Videos were obtained from authorized medical providers online. It contains the following data: A gold standard of 20 hours of 290 videos and audios, each revised by two human annotators. A silver standard of 10 hours of 76 videos and audios, in which only the medical transcripts were each revised by a single annotator. A pronunciation dictionary of medical words in Spanish, to be used with Montreal Forced Aligner. The dataset contains recordings from a total of 403 different speakers (200 male and 203 female). This repository contains only the audios and transcripts. Please, contact the author to get the corresponding videos. Acknowledgements The following linguists who contributed their time to revise and align the text and speech data of the dataset: Lara Alonso Álvaro Arozarena Miriam Lim Federico Ortega Adrián Ruiz Minnie Zheng
MedTranscripts 是一个包含30小时西班牙语医疗视频与音频的数据集,其内容已与对应字幕完成时间轴对齐。所有视频均来源于经授权的在线医疗服务提供商。 该数据集包含以下数据: 1. 金标准(gold standard)数据集:共计20小时的290条音视频素材,每条均经过两名人类标注员审核修正。 2. 银标准(silver standard)数据集:共计10小时的76条音视频素材,仅对医疗转录文本进行了审核修正。 3. 西班牙语医疗词汇发音词典,可配合蒙特利尔强制对齐器(Montreal Forced Aligner)使用。 本数据集总计收录了402位不同发言者的录音,其中男性200位、女性202位。 本存储库仅提供音频与转录文本,如需获取对应视频,请联系数据集作者。 致谢 感谢以下语言学家为本次数据集的文本与语音数据的审核及对齐工作贡献时间: Lara Alonso、Álvaro Arozarena、Miriam Lim、Federico Ortega、Adrián Ruiz、Minnie Zheng



