clean-speech-raw-sources
收藏资源简介:
该数据集是“Clean Speech”原始存档的镜像集合,旨在为原本托管在Hugging Face之外、具备干净且采样率≥44.1 kHz的语音数据集提供持久稳定的公共镜像。所有原始文件未经修改,并保留原始许可和署名。数据集包含以下三个子集:RAVDESS(来自Ryerson Audio-Visual Database of Emotional Speech and Song,仅包含语音部分,共1440个文件,24位演员,48 kHz/16-bit,许可为CC BY-NC-SA 4.0);DAPS(设备音频语音数据集,包含20位说话者,15个版本,采样率44.1 kHz,许可为CC BY-NC-SA 4.0);BibleTTS(OpenSLR 129,多语言非洲语音数据集,包含6种语言,单说话人录音室质量48 kHz,最大80小时,许可为CC BY-SA 4.0)。该镜像用于研究可复现性,训练版本位于另一个仓库“Scicom-intl/TTS-Clean44k”。
This dataset is a mirror collection of the Clean Speech original archive, intended to provide persistent and stable public mirrors of speech datasets that were originally hosted outside Hugging Face and have clean audio with sampling rate ≥44.1 kHz. All original files are unmodified, retaining original licenses and attribution. The dataset contains three subsets: RAVDESS (from Ryerson Audio-Visual Database of Emotional Speech and Song, only speech part, 1440 files, 24 actors, 48 kHz/16-bit, license CC BY-NC-SA 4.0); DAPS (Device Audio Speech Dataset, 20 speakers, 15 versions, 44.1 kHz, license CC BY-NC-SA 4.0); BibleTTS (OpenSLR 129, multilingual African speech dataset, 6 languages, single speaker studio quality 48 kHz, up to 80 hours, license CC BY-SA 4.0). This mirror is used for research reproducibility, and the training version (filtered by DNSMOS and normalized to 48 kHz) is located in another repository Scicom-intl/TTS-Clean44k.
数据集概述:Clean Speech Raw Sources (non-HF mirror)
该数据集是高质量语音原始档案的持久镜像,旨在提供不低于44.1kHz采样率的干净语音数据集的备份,并确保外部原始链接失效时仍可访问。所有档案均未修改,保留原始许可证和归属信息,仅用作镜像,原始来源链接为权威版本。
许可证:CC BY-NC-SA 4.0(非商业、相同方式共享)
主要用途:该数据集中的原始音频可作为训练数据来源,经过DNSMOS过滤和48kHz归一化处理的训练版本位于Scicom-intl/TTS-Clean44k。
数据集内容
1. RAVDESS(情感语音)
- 文件:
RAVDESS_Audio_Speech_Actors_01-24.zip - 原始来源:https://zenodo.org/records/1188976
- 音频规格:48 kHz,16位,1440个文件,24位演员(情感语音)
- 许可证:CC BY-NC-SA 4.0
- 引用:Livingstone SR, Russo FA (2018),PLoS ONE 13(5): e0196391
2. DAPS(设备与录音室语音)
- 文件:
DAPS_daps.tar.gz - 原始来源:https://zenodo.org/records/4660670
- 音频规格:44.1 kHz;20位说话人;15个版本(3个录音室版本:clean/cleanraw/produced,12个设备/房间版本)。训练时仅使用3个录音室版本。
- 许可证:CC BY-NC-SA 4.0
- 引用:Mysore GJ (2015),IEEE Signal Processing Letters 22(8)
3. BibleTTS(多语言非洲语音)
- 文件:
bibletts_openslr129/*.tgz - 原始来源:https://www.openslr.org/129/ (镜像:https://openslr.trmal.net/resources/129/)
- 音频规格:48 kHz单声道,录音室质量;每个语言最多80小时单说话人。
- 语言:Akuapem Twi、Asante Twi、Ewe、Hausa、Lingala、Yoruba(撒哈拉以南非洲语言)
- 许可证:CC BY-SA 4.0(允许自由再分发)
- 引用:Meyer et al. (2022),Interspeech 2022
备注
- 所有档案均保持原始未修改状态,许可证和署名归原作者所有。
- 该数据集仅用于研究可复现性,并作为非Hugging Face原始来源的持久镜像。




