遇见数据集

MedTranscripts - A multimodal dataset of Spanish medical videos and time-aligned transcripts

收藏
Zenodo2026-06-19 更新2026-05-26 收录
官方服务:

资源简介:

MedTranscripts is a dataset of 30 hours of medical videos and audios in Spanish, time-aligned with the corresponding transcripts. Videos were obtained from authorized medical providers online. It contains the following data: A gold standard of 20 hours of 290 videos and audios, each revised by two human annotators. A silver standard of 10 hours of 76 videos and audios, in which only the medical transcripts were each revised by a single annotator. A pronunciation dictionary of medical words in Spanish, to be used with Montreal Forced Aligner. The dataset contains recordings from a total of 403 different speakers (200 male and 203 female). This repository contains only the audios and transcripts. Please, contact the author to get the corresponding videos. Acknowledgements The following linguists who contributed their time to revise and align the text and speech data of the dataset: Lara Alonso Álvaro Arozarena Miriam Lim Federico Ortega Adrián Ruiz Minnie Zheng

提供机构:
Zenodo
创建时间:
2025-08-02
二维码
社区交流群
二维码
科研交流群
商业服务