laion/voiceclap-data
收藏资源简介:
VoiceCLAP Data数据集是一个用于训练laion/voiceclap-small和laion/voiceclap-large模型的音频和密集字幕混合数据集。每个tar分片包含配对的.flac音频文件和.json字幕及元数据文件,字幕和属性注释由音频感知LLMs自动生成,包括Qwen-Audio、Gemini Flash 2.5和EmoNet分类模型。数据集包含多个子集,如emolia(来自Emilia数据集)、laions-got-talent(来自LAIONs Got Talent数据集)、majestrino(来自Common-Voice多语言子集)等,每个子集有不同的来源和用途。数据集的语言为英语和多语言,任务类别包括音频分类和特征提取,标签涵盖音频、语音、情感等。
The VoiceCLAP Data dataset is a mixed audio and dense caption dataset designed for training the laion/voiceclap-small and laion/voiceclap-large models. Each tar shard contains paired .flac audio files alongside .json caption and metadata files. The captions and attribute annotations are automatically generated by audio-aware large language models (LLMs), including Qwen-Audio, Gemini Flash 2.5, and the EmoNet classification model. The dataset comprises multiple subsets, such as emolia (derived from the Emilia dataset), laions-got-talent (from the LAIONs Got Talent dataset), and majestrino (from the multilingual subset of Common-Voice), among others, with each subset having distinct sources and application scenarios. The dataset supports English and multiple other languages. Its task categories include audio classification and feature extraction, with labels covering audio, speech, emotion, and other relevant domains.




