common-voice-scripted-speech-25-0
收藏资源简介:
Common Voice Scripted Speech 25.0 是一个基于 Mozilla Data Collective 的 Common Voice Scripted Speech 25.0 构建的行归一化多语言自动语音识别(ASR)数据集。其主要目标是整合、规范并提供来自 Common Voice 项目的脚本化语音数据,为 ASR 研究和开发提供一个统一、可追溯的数据源。数据集的核心交付物是归一化后的数据行,每行对应一个音频片段,并保留了原始的分割信息(如训练集、开发集、测试集等)。数据字段包括音频路径、对应文本句子、语言环境代码、语言名称、数据分割类型、源数据集ID、源存档文件,以及丰富的上游元数据(如说话者ID、句子ID、领域、投票信息、人口统计信息等)。数据集规模庞大,覆盖了超过150种语言和方言(如阿布哈兹语、阿非利卡语、阿拉伯语、中文、英语、法语等),当前状态包含150个清单条目和78个已暂存的存档文件。数据集适用于多语言和低资源语言的自动语音识别任务。数据发布遵循 Creative Commons Zero v1.0 Universal (CC0-1.0) 许可证。
Common Voice Scripted Speech 25.0 is a row-normalized multilingual automatic speech recognition (ASR) dataset built upon Mozilla Data Collectives Common Voice Scripted Speech 25.0. Its primary goal is to integrate, standardize, and provide scripted speech data from the Common Voice project, offering a unified and traceable data source for ASR research and development. The core deliverable of the dataset is normalized data rows, each corresponding to an audio segment, while preserving original split information (such as training set, development set, test set, etc.). Data fields include audio paths, corresponding text sentences, locale codes, language names, data split types, source dataset IDs, source archive files, and rich upstream metadata (e.g., speaker IDs, sentence IDs, domains, voting information, demographic data, etc.). The dataset is large-scale, covering over 150 languages and dialects (e.g., Abkhaz, Afrikaans, Arabic, Chinese, English, French, etc.), with a current status of 150 manifest entries and 78 staged archive files. It is suitable for multilingual and low-resource language automatic speech recognition tasks. The data is released under the Creative Commons Zero v1.0 Universal (CC0-1.0) license.
数据集概述
数据集名称:Common Voice Scripted Speech 25.0
标签:common-voice, mozilla-data-collective, scripted-speech, multilingual-asr, asr
许可证:Creative Commons Zero v1.0 Universal (CC0-1.0),详见 CC0-1.0。
数据集描述
该数据集是一个多语言自动语音识别(ASR)数据集,基于 Mozilla Data Collective 的 Common Voice Scripted Speech 25.0 构建。主要交付物为行归一化的数据,包含音频、句子、区域设置、语言、分割、源数据集 ID、源归档等字段,并保留上游 Common Voice 元数据。
数据用途
- 提供统一的多语言 ASR 数据集,方便用户获取 Common Voice Scripted Speech 25.0 的综合集合。
- 支持通过语言和分割字段进行过滤。
- 质量过滤作为可选视图或清单,基于上游分割和验证状态,提供 Scribe WER 分级:优秀(≤5%)、良好(≤15%)、可接受(≤25%),以及持续时间和每秒字符数等回退检查。
数据集统计(当前状态)
- 清单条目数:162
- 已暂存的归档数:78
- 仍在下载或缺失的归档数:84
- 目标数据集层:归一化行,保留上游
train,dev,test,validated,invalidated,other和报告文件(如有)。
语言覆盖
数据集包含超过170种语言,涵盖多种语言家族和地区,例如:
- 亚美尼亚语 (
hy-AM) - 中文 (中国、香港、台湾) (
zh-CN,zh-HK,zh-TW) - 林加拉语 (
lg) - 荷兰语 (
nl) - 意大利语 (
it) - 等
详见完整语言列表(language 和 language_bcp47 字段)。
预期数据列
归一化表格每行对应一个音频片段,预计包含列:
audiosentencelocalelanguagesplit/original_splitsource_dataset_idsource_archive- 上游 Common Voice 元数据(如
client_id,sentence_id,sentence_domain, 投票, 年龄, 性别, 口音, 变体, 分段等) duration_ms(来自clip_durations.tsv,如可用)




