ai-coustics/dawn_chorus_en
收藏资源简介:
dawn_chorus_en 是一个开源评估数据集,用于准确的前景说话人转录。该数据集针对混合条件设计,其中前景语音通常可被语音转文本系统转录,而背景语音则明显被视为背景。它提供约90分钟的前景-背景语音混合,由录制和合成的语音组成,并包含真实的前景语音和相应的转录文本。数据集包含450对等长的mix和speech音频对,总时长为01:31:19,采样率为16kHz,单声道,16位。前景语音来源分布为65%的录制语音(19位说话人)和35%的合成语音(7位说话人)。语音性别分布为44%女性声音和56%男性声音。传输通道分布为67% GSM、16.5% WhatsApp和16.5% Telegram。数据集旨在评估在抑制背景语音的同时保留前景说话人的模型性能,适用于背景语音抑制基准测试、语音转文本鲁棒性评估、前景说话人隔离/目标说话人提取系统比较等用途。
An open-source evaluation dataset for accurate foreground speaker transcription. The dataset targets mixture conditions where foreground speech remains generally transcribable by speech-to-text systems, while background speech is distinctly perceived as background. It provides around 90 minutes of foreground–background speech mixtures composed of recorded and synthesized foreground speech, along with ground truth foreground speech and corresponding transcripts. The dataset contains 450 mix and speech pairs of equal length with a sum duration of 01:31:19, with a 16 kHz sampling rate, 16-bit, mono. Foreground speech source distribution: 65% recorded speech (19 speakers), 35% synthesized speech (7 speakers). Voice gender distribution: 44% female-sounding voices, 56% male-sounding voices. Transmission channels distribution: 67% GSM, 16.5% WhatsApp, 16.5% Telegram. It is intended for evaluation of models that suppress background speech while preserving a primary/foreground speaker in conditions relevant to downstream speech-to-text systems.




