DIHARD II
收藏资源简介:
DIHARD II数据集是由宾夕法尼亚大学语言数据联盟创建,旨在提升语音分割系统在不同录音设备、噪声环境和对话领域中的鲁棒性。该数据集包含从192个不同来源抽取的约22.49小时音频,涵盖从有声读物到儿童语言学习等多种对话场景。创建过程中,数据集经过了严格的标注和分割,确保了数据的质量和准确性。DIHARD II数据集主要用于语音分割技术的研究和开发,特别是在处理复杂交互和重叠语音方面,以期解决现有系统在特定领域或数据集上过拟合的问题。
The DIHARD II dataset was created by the Language Data Consortium at the University of Pennsylvania. It is designed to improve the robustness of speech segmentation systems across diverse recording devices, noise environments and conversational domains. The dataset contains approximately 22.49 hours of audio extracted from 192 distinct sources, covering a wide range of conversational scenarios from audiobooks to children's language learning. During its development, the dataset underwent rigorous annotation and segmentation to ensure its quality and accuracy. The DIHARD II dataset is primarily used for research and development of speech segmentation technologies, particularly in handling complex interactions and overlapping speech, with the aim of solving the overfitting issue faced by existing systems when deployed in specific domains or datasets.




