MuMu-LLaMA
收藏资源简介:
MuMu-LLaMA数据集由腾讯PCG ARC实验室和新加坡国立大学联合创建,专门用于多模态音乐理解和生成任务。该数据集包含167.69小时的文本、图像、视频和音乐注释,通过先进的视觉模型如LLaVA和Video-LLaVA进行标注,确保数据的多样性和高质量。数据集的创建过程结合了多种模态的特征提取和标注技术,旨在为多模态音乐研究提供丰富的训练数据。该数据集主要应用于音乐理解和生成领域,旨在解决多模态输入下的音乐理解和生成问题,推动音乐创作和分析的创新应用。
The MuMu-LLaMA dataset was jointly developed by Tencent PCG ARC Lab and the National University of Singapore, and is specifically tailored for multimodal music understanding and generation tasks. This dataset contains 167.69 hours of text, image, video and music annotations, which are annotated via state-of-the-art vision-language models including LLaVA and Video-LLaVA to guarantee the diversity and high quality of the dataset. The development pipeline of the dataset integrates multimodal feature extraction and annotation techniques, aiming to provide abundant training data for multimodal music research. Primarily deployed in the domains of music understanding and generation, this dataset aims to address the challenges of music understanding and generation under multimodal inputs, and promote innovative applications of music creation and analysis.




