yodas_owsmv4
收藏资源简介:
该数据集包含了跨越75种语言的166,000小时的多种语言语音,被分割成30秒的长格式音频片段。数据来源于YODAS2数据集,该数据集基于大规模的网络爬取内容。由于网络源数据的性质,原始的YODAS2数据集可能包含不准确的语言标签和音频-文本对不齐的情况。为了解决这一问题,我们开发了一个可扩展的数据清洗管道,使用公开可用的工具包,从而形成原始数据集的一个精选子集。这个清洗后的数据集是我们OWSM v4模型训练数据的核心部分,结合现有的OWSM数据,这些模型在多种语言自动语音识别基准测试中的表现显著优于以前版本。
This dataset contains 166,000 hours of multilingual speech spanning 75 languages, segmented into 30-second long-form audio clips. It is derived from the YODAS2 dataset, which is built on large-scale web-crawled content. Due to the nature of web-sourced data, the original YODAS2 dataset may contain inaccurate language labels and misaligned audio-text pairs. To address these issues, we developed a scalable data cleaning pipeline using publicly available toolkits to create a curated subset of the original dataset. This cleaned dataset forms a core component of the training data for our OWSM v4 models; when combined with existing OWSM data, these models achieve significantly better performance on multilingual automatic speech recognition benchmarks than previous versions.




