Hien141625/Dataset_STT
收藏资源简介:
Dolly-Audio是一个大规模、高质量的越南语语音语料库,由Dolly AI Team创建。灵感来自世界上第一个克隆哺乳动物多莉,该项目旨在推动越南语语音合成、语音识别和语音建模的研究。该版本提供近1,000小时的专业清理音频,涵盖152名来自不同越南地区和说话风格的说话者。文本转录跨越多个领域,以确保语言多样性和模型鲁棒性。关键特征包括:高质量越南语语音约1,000小时、152名多地区说话者带不同口音、清理无噪声的音频无背景音乐、句子级别边界修剪以自然韵律、丰富的转录领域(如新闻、娱乐、教育、对话等)、通过手动采样估计词错误率接近零(≈0%)。数据集适用于多说话者文本到语音、自动语音识别、语音克隆和说话者适应、韵律建模、语言和语音学研究。使用限制:仅限非商业研究用途,必须遵守CC-BY-NC-SA-4.0许可。
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team. Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling. This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and speaking styles. Text transcripts span a wide variety of domains to ensure linguistic diversity and model robustness. Key features include: ~1,000 hours of high-quality Vietnamese speech, 152 multi-region speakers with diverse accents, cleaned, noise-free audio with no background music, sentence-level boundary trimming for natural prosody, rich transcript domains (news, entertainment, education, conversational, etc.), estimated near-zero WER (≈ 0%) from manual sampling. The dataset is suitable for multi-speaker text-to-speech, automatic speech recognition, voice cloning and speaker adaptation, prosody modeling, and linguistic and phonetic research. Usage restrictions: non-commercial research use only, must comply with CC-BY-NC-SA-4.0.




