Cong123779/data
收藏资源简介:
这是一个大规模的越南语文本到语音(TTS)数据集,包含来自网络文学和翻译故事的高质量音频及相应文本,总容量为118GB。数据集旨在支持现代TTS模型的训练,如Matcha-TTS、F5-TTS和Piper。数据集分为三个主要部分:1) Thế Giới Hoàn Mỹ(武侠风格,适合小说TTS),2) Án Sát(侦探题材,多样化的角色对话),3) Ngạo Thế Cửu Trọng Thiên(已降噪并标准化为22050Hz,可直接用于训练)。音频格式为22k-44k Hz的.wav文件(单声道/立体声),元数据为.csv格式。数据集总时长超过1000小时。由于Hugging Face的50GB限制,大文件被分割成多个部分,需使用提供的命令进行合并和解压。
This is a large-scale Vietnamese Text-to-Speech (TTS) dataset, comprising high-quality audio and corresponding texts from online literature and translated stories, totaling 118GB. The dataset is designed to facilitate the training of modern TTS models such as Matcha-TTS, F5-TTS, and Piper. It is divided into three main projects: 1) Thế Giới Hoàn Mỹ (martial arts style, suitable for novel TTS), 2) Án Sát (detective theme, featuring diverse character dialogues), and 3) Ngạo Thế Cửu Trọng Thiên (noise-filtered and standardized to 22050Hz, ready for immediate training). The audio is in .wav format (mono/stereo) at 22k-44k Hz, and metadata is in .csv format. The total duration exceeds 1000 hours. Due to Hugging Faces 50GB limit, large files are split into parts and require merging and extraction using provided commands.



