Dataset_STT
收藏资源简介:
Dolly-Audio是一个大规模、高质量的越南语语音语料库,由Dolly AI团队创建。该数据集旨在推进越南语语音合成、语音识别和语音建模的研究。数据集包含近1,000小时经过专业清洗的音频数据,涵盖152位来自越南不同地区、具有不同口音和说话风格的说话人。文本转录本覆盖新闻、娱乐、教育、对话等多种领域,确保语言多样性和模型鲁棒性。音频经过清洗,无背景噪音和音乐,并进行了句子级边界修剪以获得自然的韵律。数据集包含664,125个样本,每个样本包含音频文件名、文本转录、说话人ID和音频数据四个字段。该数据集适用于多说话人文本到语音合成、自动语音识别、语音克隆与说话人适应、韵律建模以及语言学和语音学研究。数据集仅限非商业研究使用,遵循CC-BY-NC-SA-4.0许可协议。
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI team. This dataset aims to advance research in Vietnamese speech synthesis, automatic speech recognition (ASR), and speech modeling. The dataset contains nearly 1,000 hours of professionally cleaned audio data, covering 152 speakers from different regions of Vietnam with diverse accents and speaking styles. Its text transcripts cover multiple domains including news, entertainment, education, dialogue, etc., ensuring linguistic diversity and model robustness. The audio has been cleaned to remove background noise and music, and trimmed at sentence-level boundaries to achieve natural prosody. The dataset consists of 664,125 samples, each containing four fields: audio filename, text transcript, speaker ID, and audio data. This dataset is suitable for multi-speaker text-to-speech (TTS) synthesis, automatic speech recognition, voice cloning and speaker adaptation, prosody modeling, as well as linguistic and phonetic research. The dataset is for non-commercial research use only and follows the CC-BY-NC-SA-4.0 license agreement.
数据集概述
数据集名称:Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
发布方:Dolly AI Team
语言:越南语 (vi)
标签:vietnamese, synthetic, audio, tts
数据集规模:约10万至100万条样本(100K < n < 1M)
数据文件:训练集(train)共664,125个样本,数据集总大小约165.96 GB,下载大小约157.80 GB
数据特征:
- audio_filename(字符串)
- text(字符串)
- voice_id(字符串)
- audio(音频类型,未解码)
核心特性:
- 约1,000小时高质量越南语语音
- 152位跨区域说话人,涵盖多种口音
- 音频经过专业清洁,无噪声、无背景音乐
- 句子级边界修剪,自然韵律保留
- 文本转录涵盖新闻、娱乐、教育、对话等多领域
- 手动抽样评估接近零词错误率(≈0%)
- 适用于文本转语音(TTS)、自动语音识别(ASR)、语音克隆和语音研究
预期用途:
- 多说话人文本转语音(TTS)
- 自动语音识别(ASR)
- 语音克隆与说话人自适应
- 韵律建模
- 语言学与语音学研究
使用限制:
- 仅限非商业研究用途
- 再分发须遵守CC-BY-NC-SA-4.0许可
- 用户需自行验证数据集是否适合研究任务
- 访问审批需使用机构邮箱
引用格式(BibTeX):
@dataset{dolly_audio_2025, title = {Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus}, author = {Nguyen, Vinh Huy and Nguyen, Dinh Thuan}, year = {2025}, publisher = {Dolly AI Team}, howpublished = {url{https://huggingface.co/datasets/Dolly-AI/Dolly-Audio}}, note = {Released under CC-BY-NC-SA-4.0. Research use only.} }
联系方式:Nguyen Vinh Huy(nguyenvinhhuy@dtu.edu.vn)及 Nguyen Dinh Thuan(boyphuthien115@gmail.com)




