Muse Synthetic Song Dataset
收藏资源简介:
Muse合成歌曲数据集由复旦大学自然语言处理组构建,包含11.6万首全授权合成歌曲,总时长约7,771小时。该数据集通过GPT生成歌词与风格标签,经SunoV5合成音频,并采用Qwen3-Omni模型进行全局和段落级风格标注,平均每首歌曲时长为4分钟。数据涵盖中英文双语言(中文4.9万/英文6.6万首),支持长格式歌曲生成研究,解决了学术领域可复现数据匮乏的痛点,适用于音乐生成模型的训练与风格可控性研究。
The Muse synthetic song dataset was constructed by the Natural Language Processing Group of Fudan University. It contains 116,000 fully authorized synthetic songs with a total duration of approximately 7,771 hours. The dataset generates lyrics and style tags via GPT, synthesizes audio using SunoV5, and performs global and paragraph-level style annotation with the Qwen3-Omni model. The average duration of each song is 4 minutes. The data covers both Chinese and English languages, with 49,000 Chinese songs and 66,000 English songs respectively. It supports research on long-form song generation, addressing the pain point of scarce reproducible data in the academic field, and is suitable for training music generation models and research on style controllability.
Muse数据集概述
数据集名称
Muse
数据集简介
Muse是一个面向可复现长歌曲生成的数据集,其核心目标在于实现细粒度的风格控制。




