遇见数据集

Alright7398/modded-distill-wavlm-base

收藏
Hugging Face2026-05-05 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含LibriSpeech语料库在16 kHz采样率下的音频数据,以及通过冻结的microsoft/wavlm-base模型预计算的最后一层隐藏状态表示。这是一个预计算的教师缓存和原始音频字节数据集,并非通用的语音基准测试分割。数据集结构包括训练集(train/)和评估集(eval/),分别对应LibriSpeech的960小时训练语料(train-clean-100、train-clean-360、train-other-500)和开发集(dev-clean)。每个数据行包含音频键、分割标签、原始音频字节、音频长度(以样本数表示)以及教师模型的最后一层隐藏状态(以float16列表形式存储)。数据集适用于特征提取和表示学习任务,使用Lance格式存储,需通过Lance读取器加载。

Lance tables of LibriSpeech utterances at 16 kHz with cached last-layer representations from frozen microsoft/wavlm-base. This release is a precomputed teacher cache plus raw audio bytes, not a general-purpose speech benchmark split. The dataset includes train and eval directories, corresponding to the LibriSpeech 960-hour training corpora (train-clean-100, train-clean-360, train-other-500) and dev-clean set, respectively. Each row contains an utterance key, split tag, raw audio bytes, audio length in samples, and flattened last hidden states from WavLM-Base for valid frames (stored as float16 list). It is designed for feature extraction and representation learning, stored in Lance format and requires a Lance reader for loading.

提供机构:
Alright7398
二维码
社区交流群
二维码
科研交流群
商业服务