ConceptCaps
收藏资源简介:
ConceptCaps是由华沙理工大学团队构建的面向音乐模型可解释性研究的高质量数据集,包含23,000条音乐-文本-音频三元组数据。该数据集基于200个音乐属性标签体系构建,采用三阶段生成流程:首先通过变分自编码器建模音乐属性的共现模式,再由微调的大语言模型生成专业描述,最后通过MusicGen合成版权自由的对应音频。数据内容涵盖乐器、流派、情绪等多维度音乐概念,其结构化标签设计有效解决了传统音乐数据标签稀疏、噪声大的问题,特别适用于TCAV等基于概念的模型可解释性分析方法。
ConceptCaps is a high-quality dataset developed by the team from Warsaw University of Technology for research on the interpretability of music models, containing 23,000 music-text-audio triplet samples. Built upon a taxonomy of 200 music attribute labels, this dataset follows a three-stage generation pipeline: first, it models the co-occurrence patterns of music attributes using a Variational Autoencoder (VAE); second, it generates professional descriptions via a fine-tuned Large Language Model (LLM); finally, it synthesizes copyright-free corresponding audio through MusicGen. Covering multi-dimensional music concepts including musical instruments, genres, emotions and more, its structured label design effectively addresses the issues of sparse labeling and high noise in traditional music data, making it particularly suitable for concept-based model interpretability analysis methods such as TCAV.



