ASLP-lab/UrduSpeech
收藏资源简介:
UrduSpeech是一个大规模、高保真的乌尔都语语音语料库,包含156小时的音频,并附有全面的12维副语言元数据。该语料库通过以下方式解决了乌尔都语在语音技术中资源严重不足的问题:- 71,792个经过对话者分割的话语,涵盖多样化的内容类别;- 三个专业子集:标准巴基斯坦乌尔都语(US-Std,59.2小时)、乌尔都语-英语语码转换(US-CS,89.4小时)和巴基斯坦口音英语(US-EngPk,7.3小时);- 12个内容类别:喜剧节目、戏剧、电影、食品、访谈、新闻、播客、诗歌、散文、路边采访、视频博客、YouTube评论;- 丰富的副语言注释:性别、年龄、音高、语速、情感、口音、语调、节奏、音质、发音、副语言特征和上下文信息;- 高质量验证:平均意见分数(MOS)为4.64(σ=0.7),评分者间一致性Cohens Kappa为0.68;- 性别平衡:话语中男女比例为60/40;- 转录置信度:模型生成并经过人工验证的转录置信度为97.6%。该语料库采用严格的LLM驱动流程(使用Gemini 2.5 Pro)进行整理,解决了乌尔都语的独特挑战,包括从右到左(RTL)脚本限制、乌尔都语-英语语码转换以及与印地语的声学接近性。单独发布的9小时手动校正基准集(US-benchmark)作为评估的黄金标准。
UrduSpeech is a large-scale, high-fidelity Urdu speech corpus comprising 156 hours of audio with comprehensive 12-dimensional paralinguistic metadata. The corpus addresses the critical under-resourcing of Urdu in speech technology by providing: - 71,792 diarized utterances across diverse content categories; - Three specialized subsets: Standard Pakistani Urdu (US-Std, 59.2h), Urdu-English Code-Switched (US-CS, 89.4h), and Pakistani-Accented English (US-EngPk, 7.3h); - 12 content categories: Comedy Show, Drama, Film, Food, Interview, News, Podcast, Poetry, Proses, Roadside Interview, Vlogs, YouTube Review; - Rich paralinguistic annotations: gender, age, pitch, speed, emotion, accent, tone, rhythm, texture, pronunciation, paralinguistic features, and contextual information; - High-quality validation: Mean Opinion Score (MOS) of 4.64 (σ = 0.7) with 0.68 Cohens Kappa inter-rater reliability; - Gender balance: 60/40 distribution across utterances; - Transcription confidence: 97.6% confidence score with model-generated and manually-validated transcriptions. The corpus was curated using a rigorous LLM-driven pipeline with Gemini 2.5 Pro, addressing Urdus unique challenges including Right-to-Left (RTL) script constraints, Urdu-English code-switching, and acoustic proximity to Hindi. A separately released 9-hour manually-corrected benchmark set (US-benchmark) serves as the gold standard for evaluation.




