遇见数据集

ryota-komatsu/SylReg-Chunk1

收藏
Hugging Face2026-05-15 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个多配置集合,包含Hi-Fi-CAPTAIN、Libri-Light、LibriTTS-R、sBLIMP、sSC、sWUGGY和tSC等子数据集,主要用于语音合成、语言建模和语音处理任务。数据集特征包括音频单元(units)、持续时间(durations)、文本转录(transcript)和声谱图(spectrogram)等,支持训练、验证和测试分割。例如,Hi-Fi-CAPTAIN包含女性语音数据,Libri-Light和LibriTTS-R基于大型语音语料库,sBLIMP、sSC、sWUGGY和tSC则涉及对比语言结构分析。数据以结构化格式存储,适用于机器学习和NLP研究。

This dataset is a multi-configuration collection including sub-datasets such as Hi-Fi-CAPTAIN, Libri-Light, LibriTTS-R, sBLIMP, sSC, sWUGGY, and tSC, primarily designed for speech synthesis, language modeling, and speech processing tasks. The features include audio units, durations, text transcripts, and spectrograms, with support for train, validation, and test splits. For instance, Hi-Fi-CAPTAIN contains female speech data, Libri-Light and LibriTTS-R are based on large speech corpora, while sBLIMP, sSC, sWUGGY, and tSC involve contrastive linguistic structure analysis. The data is stored in a structured format suitable for machine learning and NLP research.

提供机构:
ryota-komatsu
二维码
社区交流群
二维码
科研交流群
商业服务