KRAFTON/Raon-OpenTTS-Pool
收藏资源简介:
Raon-OpenTTS-Pool是一个用于文本到语音(TTS)训练的大规模开放英语语音语料库,由8个公开可用的语音语料库和一组网络来源的录音构建而成。它是Raon-OpenTTS训练数据的基础,Raon-OpenTTS是一个开放TTS模型,其性能与最先进的封闭数据系统相当。该数据集包含615K小时的语音音频、239.7M个语音片段,所有音频以16 kHz单声道Opus(64 kbps)格式存储在WebDataset tar分片中。数据集还包括一个经过质量过滤的子集Raon-OpenTTS-Core,包含510.1K小时和194.5M个片段。数据来源限制为公开可用的英语语音数据集,每个音频片段限制在30秒或更短,以减少对齐错误、多说话人内容和非语音伪影。现有公共数据集(如LibriHeavy、Emilia、VoxPopuli等)未经修改直接包含,音频标准化为16 kHz单声道Opus 64 kbps以提高存储效率。Raon-YouTube-Commons部分通过专用预处理流程从YouTube-Commons重建。
Raon-OpenTTS-Pool is a large-scale open English speech corpus for text-to-speech (TTS) training, constructed from 8 publicly available speech corpora and a set of web-sourced recordings. It serves as the foundational training data for Raon-OpenTTS, an open-source TTS model whose performance is on par with state-of-the-art closed-data systems. This dataset contains 615K hours of speech audio and 239.7M speech segments, with all audio stored in WebDataset tar shards in 16 kHz mono Opus (64 kbps) format. The dataset also includes a quality-filtered subset named Raon-OpenTTS-Core, which consists of 510.1K hours and 194.5M segments. All data sources are restricted to publicly available English speech datasets, and each audio segment is limited to 30 seconds or shorter to reduce alignment errors, multi-speaker content, and non-speech artifacts. Existing public datasets such as LibriHeavy, Emilia, VoxPopuli and others are included unmodified, and audio is standardized to 16 kHz mono Opus 64 kbps to improve storage efficiency. The Raon-YouTube-Commons subset is reconstructed from YouTube-Commons via a dedicated preprocessing pipeline.




