UrduSpeech
收藏资源简介:
UrduSpeech是一个大规模、高保真的乌尔都语语音语料库,总时长156小时,包含71,792个经过说话人分割的话语,并附有全面的12维副语言学元数据。该语料库旨在解决乌尔都语在语音技术领域资源严重不足的问题。数据集覆盖12个内容类别:喜剧节目、戏剧、电影、美食、访谈、新闻、播客、诗歌、散文、街头采访、视频博客和YouTube评论。它由三个专门的子集构成:标准巴基斯坦乌尔都语(US-Std,59.2小时)、乌尔都语-英语语码转换(US-CS,89.4小时)和巴基斯坦口音英语(US-EngPk,7.3小时)。每个子集进一步按音频时长分为短片段(≤10秒,共55,407段)和长片段(>10-35秒,共16,243段)。数据收集自YouTube和巴基斯坦电视台(PTV)的档案内容,时间跨度超过40年(从1980年代至今),覆盖巴基斯坦及海外巴基斯坦人社区,确保了说话人(包括媒体专业人士和非专业人士)和口音的多样性。每个音频实例都配有详细的元数据,包括:转录文本(具有97.6%的置信度评分)、说话人ID、音频类别、时长、词数、字符数以及12维的副语言学标注(如性别、年龄、音高、语速、情感、口音、语调、节奏、音质、发音、副语言学特征和上下文信息)。数据集在性别分布上保持平衡(男性60%,女性40%),并包含不同年龄段的说话人。质量评估显示其平均意见得分(MOS)为4.64(σ=0.7),评分者间一致性Cohens Kappa为0.68。此外,还单独发布了一个经过人工全面校正的9小时基准集(US-benchmark),用于评估。该数据集适用于多种语音任务,包括自动语音识别(ASR)、语码转换研究、语音情感识别、说话人画像(年龄、性别、口音分类)、副语言学分析、文本到语音(TTS)、语音增强以及多语言语音模型训练。数据集采用Creative Commons Attribution 4.0 International (CC-BY-4.0) 许可证发布。
UrduSpeech is a large-scale, high-fidelity Urdu speech corpus with a total duration of 156 hours, containing 71,792 speaker-segmented utterances, and accompanied by comprehensive 12-dimensional paralinguistic metadata. The corpus aims to address the severe shortage of resources for Urdu in speech technology. The dataset covers 12 content categories: comedy shows, drama, movies, food, interviews, news, podcasts, poetry, prose, street interviews, vlogs, and YouTube comments. It consists of three specialized subsets: Standard Pakistani Urdu (US-Std, 59.2 hours), Urdu-English code-switching (US-CS, 89.4 hours), and Pakistani-accented English (US-EngPk, 7.3 hours). Each subset is further divided by audio duration into short segments (≤10 seconds, totaling 55,407 segments) and long segments (>10-35 seconds, totaling 16,243 segments). Data was collected from YouTube and Pakistan Television (PTV) archives, spanning over 40 years (from the 1980s to present), covering Pakistan and overseas Pakistani communities, ensuring diversity in speakers (including media professionals and non-professionals) and accents. Each audio instance is accompanied by detailed metadata, including: transcription text (with a 97.6% confidence score), speaker ID, audio category, duration, word count, character count, and 12-dimensional paralinguistic annotations (such as gender, age, pitch, speech rate, emotion, accent, intonation, rhythm, voice quality, pronunciation, paralinguistic features, and context). The dataset maintains balanced gender distribution (60% male, 40% female) and includes speakers from different age groups. Quality assessment shows a Mean Opinion Score (MOS) of 4.64 (σ=0.7) and inter-rater agreement Cohens Kappa of 0.68. Additionally, a manually corrected 9-hour benchmark set (US-benchmark) is separately released for evaluation. The dataset is suitable for various speech tasks, including automatic speech recognition (ASR), code-switching research, speech emotion recognition, speaker profiling (age, gender, accent classification), paralinguistic analysis, text-to-speech (TTS), speech enhancement, and multilingual speech model training. The dataset is released under the Creative Commons Attribution 4.0 International (CC-BY-4.0) license.
数据集概括
- 数据集名称: UrduSpeech
- 数据集规模: 156 小时高质量乌尔都语语音,包含 71,792 个已进行说话人分离的语音片段。
- 质量评估: 平均意见得分 (MOS) 为 4.64 (标准差 0.7),Cohen’s Kappa 评分者间信度为 0.68,转录置信度高达 97.6%。
- 许可协议: Creative Commons Attribution 4.0 International (CC-BY-4.0)。
数据集子集
数据集包含三个专业子集:
- US-Std (标准巴基斯坦乌尔都语): 59.2 小时
- US-CS (乌尔都语-英语代码混合): 89.4 小时
- US-EngPk (巴基斯坦口音英语): 7.3 小时
数据结构与实例
数据集根目录下按子集组织,每个子集分为 short (≤10 秒,共 55,407 段) 和 long (>10 秒,共 16,243 段) 两个文件夹。每个文件夹下按 12 种内容类别(喜剧节目、戏剧、电影、美食、访谈、新闻、播客、诗歌、散文、街头采访、Vlog、YouTube 评论)存放音频文件及对应的 JSONL 元数据文件。
每条数据实例包含两类 JSONL 文件:
- 转录文件: 提供说话人 ID、音频类别、时长、词数、字符数、转写文本、置信度分数等信息。
- 副语言特征文件: 提供说话人的性别、年龄、音高、语速、情感、口音、语调、节奏、音色、发音、副语言特征(如背景音)和语境信息等多达 12 维的元数据。
数据统计
- 说话人性别分布: 女性 40% (28,802 段),男性 60% (42,990 段)。
- 说话人年龄分布: 青年 (34,126 段)、中年 (33,495 段)、儿童 (1,804 段)、老年 (2,367 段)。
- 唯一说话人数量: 估计超过 1,000 人。
数据来源与收集
数据来源于 YouTube 和巴基斯坦电视台 (PTV) 自 1980 年代至今的档案内容,涵盖媒体专业人士和非专业发言人(如Vlog博主、街头采访者、海外巴基斯坦人),地理覆盖巴基斯坦及巴基斯坦侨民。
预期用途
- 自动语音识别 (ASR)
- 代码混合语言研究
- 语音情感识别
- 说话人画像(年龄、性别、口音分类)
- 副语言特征分析
- 文本转语音 (TTS)
- 语音增强
- 多语言语音模型训练
非预期用途
- 未经同意识别或追踪个人
- 生成模仿特定个人的合成语音
- 违反隐私或文化规范的内容
基准测试
附带一个独立的 US-benchmark 基准测试集(9 小时),所有转录由母语者人工校对,修正了代码混合歧义,并保留了 12 维副语言元数据,按三个子集、12 种内容类别及长短音频格式进行划分。
附加资源
- Demo 演示: https://interspeech-urdu-demo.github.io/corpus-demo/
- 论文: https://arxiv.org/abs/2605.17846




