corti/med-term
收藏资源简介:
Med-Term是一个自动语音识别(ASR)评估数据集,由Corti ApS发布,伴随《Symphony for Speech Recognition》白皮书。该数据集包含完全合成的医疗笔记,使用文本转语音(TTS)技术生成,涵盖德语、法语和英语三种语言。每个语言配置(de、fr、en)均有200个测试样本,总计600个样本。数据不包含真实患者信息、受保护健康信息(PHI)或可识别的第三方内容。特征包括音频(采样率44100 Hz,单声道)、转录文本、持续时间、听写命令、性别、说话者ID、格式化转录、医学术语及其格式化版本、格式化术语等。数据集旨在用于评估和基准测试ASR及相关NLP系统、学术研究,以及复现白皮书结果,但禁止用于训练模型、声音模仿、竞争性产品开发、深度伪造、临床决策或重新识别个体。许可证基于CDLA-Permissive-2.0和Corti使用限制附录。
Med-Term is an evaluation dataset released by Corti ApS alongside the *Symphony for Speech Recognition* white-paper. It is a fully synthetic medical notes dataset dictated using TTS in German, French, and English, built for benchmarking automatic speech recognition (ASR) and related NLP systems on medical-domain audio. The dataset contains no real patient data, PHI, or identifiable third-party content. It includes three language configurations (de, fr, en), each with 200 test examples, totaling 600 samples. Features include audio (sampling rate 44100 Hz, mono), transcription, duration, dictation command, gender, speaker_id, transcription_formatted, medical_terms, medical_terms_formatted, and formatting_terms. Intended uses are evaluating ASR/NLP systems, academic research, and reproducing white-paper results, but it is restricted from training models, voice imitation, competitive product development, deepfakes, clinical decisions, or re-identification. Licensed under CDLA-Permissive-2.0 with Corti Use Restrictions Addendum.




