遇见数据集

thewh1teagle/snac-audio-datasets

收藏
Hugging Face2026-05-11 更新2026-05-31 收录
官方服务:

资源简介:

LibriHeavy SNAC是一个基于LibriHeavy大型分割版本的数据集,经过文本过滤处理,并使用SNAC音频令牌进行编码。数据集包含主数据集和说话人参考集:主数据集有10,930,272行,总时长为44,634.32小时,列包括ID、说话人ID、书籍ID、音频时长、源音频路径、源采样率、原始文本、转录文本、SNAC令牌(snac_0、snac_1、snac_2)和音素;SNAC元数据基于模型hubertsiuzdak/snac_24khz,采样率为24000,代码类型为uint16;文本过滤基于原始文本列,保留包含ASCII字母、数字、标点、空格、短破折号和长破折号的行;音素处理使用espeak后端,语言为美式英语,保留重音和标点。说话人参考集位于speakers/路径下,有30,132行,涉及6,190个说话人,总时长为89.01小时,每个说话人最多有5个参考片段,音频为原始Ogg/Opus格式。该数据集主要用于文本到语音和音频到音频任务,语言为英语。

Filtered LibriHeavy large split encoded with SNAC audio tokens. The dataset includes a main dataset and a speaker reference set: the main dataset has 10,930,272 rows with a total duration of 44,634.32 hours, columns include id, speaker_id, librivox_book_id, audio_duration, source_audio_path, source_sample_rate, text_original, text_transcription, snac_0, snac_1, snac_2, and phonemes; SNAC metadata is based on the model hubertsiuzdak/snac_24khz with a sample rate of 24000 and code dtype uint16; text filtering is applied to the text_original column, keeping rows containing ASCII letters, digits, punctuation, whitespace, en dash, and em dash; phonemes are processed using the espeak backend with en-us language, stress enabled, and punctuation preserved. The speaker reference set is located at speakers/ with 30,132 rows, 6,190 speakers, and a total duration of 89.01 hours, each speaker having up to 5 reference clips with audio in original Ogg/Opus format. The dataset is intended for text-to-speech and audio-to-audio tasks, with English as the primary language.

提供机构:
thewh1teagle
二维码
社区交流群
二维码
科研交流群
商业服务