voxlingua107_wds
收藏资源简介:
VoxLingua107是一个用于训练口语语言识别模型的语音数据集。该数据集由从YouTube视频自动提取的短语音片段组成,并根据视频标题和描述的语言进行标记,经过一些后处理步骤过滤掉假阳性。VoxLingua107包含107种语言的数据,训练集中的语音总量为6628小时,平均每种语言的数据量为62小时,但实际每种语言的数据量差异较大。还有一个独立的发展集,包含来自33种语言的1609个语音片段,由至少两名志愿者验证确实包含给定的语言。
VoxLingua107 is a speech dataset designed for training spoken language identification models. This dataset is composed of short speech segments automatically extracted from YouTube videos, labeled based on the languages of the video titles and descriptions, with false positives filtered out via several post-processing steps. VoxLingua107 covers data for 107 languages in total. The total duration of speech in the training set amounts to 6628 hours, with an average of 62 hours per language, yet the actual data volume varies significantly across different languages. There is also a separate development set containing 1609 speech segments from 33 languages, which have been verified by at least two volunteers to indeed contain the given language.
VoxLingua107 数据集概述
数据集简介
VoxLingua107 是一个用于训练口语语言识别模型的语音数据集。该数据集包含从 YouTube 视频中自动提取的短语音片段,并根据视频标题和描述的语言进行标注,经过后处理步骤过滤误报。
关键特性
- 语言数量:包含 107 种语言的数据
- 总数据量:训练集语音总时长为 6628 小时
- 平均数据量:每种语言平均 62 小时(实际各语言数据量差异较大)
- 开发集:包含来自 33 种语言的 1609 个语音片段,经至少两名志愿者验证确认为对应语言
数据收集方法
通过使用语言特定搜索短语(从各语言维基百科随机选取)检索 YouTube 视频,提取音频数据。若视频标题和描述的语言与搜索短语语言匹配,则认为该视频音频可能为对应语言。采用语音/非语音检测和说话人日志技术将视频分割为短句级语音片段,并通过数据驱动的后过滤步骤移除与同语言其他片段差异过大的片段(可能非目标语言)。
数据质量说明
由于自动数据收集过程,数据集中仍存在约 2% 的片段非目标语言或包含非语音内容,某些语言(如威尔士语)此比例较高。
许可证信息
本数据集基于 Creative Commons Attribution 4.0 International License 分发。版权归视频原始所有者所有。
偏差说明
数据集中语言、口音、方言、性别、种族和社会因素的分布不能代表全球人口分布。使用此数据集训练和部署模型可能引入意外偏差。
语言数据量详情
数据集涵盖 107 种语言,各语言数据量详见原始表格(包含语言代码、语言名称和小时数)。
引用信息
如需引用本数据集,请使用以下格式:
@inproceedings{valk2021slt, title={{VoxLingua107}: a Dataset for Spoken Language Recognition}, author={J{"o}rgen Valk and Tanel Alum{"a}e}, booktitle={Proc. IEEE SLT Workshop}, year={2021}, }




