MoDiCoL
收藏资源简介:
MoDiCoL是由汉堡大学知识技术实验室创建的模块化诊断持续学习语音数据集,旨在系统评估自动语音识别模型在受控环境下的鲁棒性演化。该数据集包含8100个样本,总计18.79小时语音数据,其中14.08小时为合成语音,平均样本时长为8.35秒,数据来源于LibriSpeech、Common Voice等开源语音数据集及医学、航空管制等领域的专业文本语料。数据集通过正交实验设计方法构建,涵盖语言内容、说话者特征和声学环境三大类九个子因素,并采用合成语音生成与语音增强流水线技术填补现实数据缺口。该数据集主要应用于语音识别模型的持续学习研究,旨在解决模型在真实场景中因录音条件、口音、语音障碍等多因素复合分布偏移导致的性能退化问题,为分析鲁棒性的获取、迁移与遗忘机制提供标准化测试平台。
MoDiCoL is a modular diagnostic continual learning speech dataset created by the Knowledge Technology Lab at the University of Hamburg, aiming to systematically evaluate the robustness evolution of automatic speech recognition (ASR) models in controlled environments. This dataset contains 8100 samples with a total of 18.79 hours of speech data, among which 14.08 hours are synthetic speech, and the average sample duration is 8.35 seconds. The data is sourced from open-source speech datasets such as LibriSpeech and Common Voice, as well as professional text corpora from fields like medicine and air traffic control. The dataset is constructed via orthogonal experimental design, covering three major categories and nine sub-factors including linguistic content, speaker characteristics, and acoustic environment. It adopts synthetic speech generation and speech enhancement pipelines to fill the gaps in real-world data. This dataset is mainly applied to continual learning research on speech recognition models, aiming to address the performance degradation caused by compound distribution shifts of multiple factors such as recording conditions, accents, and speech impairments in real-world scenarios, providing a standardized testbed for analyzing the acquisition, transfer, and forgetting mechanisms of robustness.
数据集概述:MoDiCoL
- 数据集名称:MoDiCoL: A Modular Diagnostic Continual Learning Dataset for ASR
- 许可证:Creative Commons Attribution 4.0 International (CC-BY-4.0)
- 任务类型:
- 自动语音识别
- 音频分类
- 语言:英语
- 数据集规模:样本数量在 1,000 到 10,000 之间
- 数据集简介:该数据集是一个为自动语音识别(ASR)设计的模块化诊断持续学习数据集。




