The People’s Speech
收藏资源简介:
The People’s Speech是一个大规模、多样化的英语语音识别数据集,由哈佛大学等研究机构创建,包含30,000小时的语音数据,主要来源于互联网档案馆。数据集通过搜索互联网上具有适当许可的音频数据及其现有转录来收集,使用Apache 2.0许可发布其数据收集系统。该数据集旨在解决自动语音识别系统训练数据的质量和多样性问题,特别是在商业应用中,通过提供大量、多样化的语音数据来提高模型的准确性和泛化能力。
The People’s Speech is a large-scale, diverse English automatic speech recognition dataset created by research institutions including Harvard University. It contains 30,000 hours of speech data primarily sourced from the Internet Archive. The dataset is collected by searching for properly licensed audio data and their existing transcriptions on the Internet. Its data collection system is released under the Apache 2.0 license. This dataset aims to address the issues of quality and diversity in training data for automatic speech recognition systems, especially in commercial applications, by providing large volumes of diverse speech data to improve the accuracy and generalization ability of the models.




