MulTTiPop
收藏资源简介:
MulTTiPop是由卡内基梅隆大学创建的流行音乐多轨转录数据集,旨在为自动音乐转录模型评估提供基准。该数据集包含572个音乐片段,总计3.5小时音频,源自1930年代至2000年代多样流派和年代的流行歌曲,数据来源于Lakh MIDI和TheoryTab数据集。其创建过程通过元数据匹配音频与MIDI文件,结合节拍追踪和人工锚定节拍选择,实现精确的时间对齐。该数据集主要应用于自动音乐转录领域,旨在解决商业流行音乐多轨转录评估数据缺乏的问题,推动模型在真实世界音乐场景中的性能提升。
MulTTiPop is a multi-track transcription dataset of popular music developed by Carnegie Mellon University, serving as a benchmark for evaluating automatic music transcription models. This dataset comprises 572 music clips totaling 3.5 hours of audio, sourced from pop songs across diverse genres and eras spanning the 1930s to the 2000s, and is derived from the Lakh MIDI and TheoryTab datasets. Its creation workflow matches audio and MIDI files via metadata, combined with beat tracking and manual anchor beat selection to achieve precise temporal alignment. Primarily applied in the field of automatic music transcription, this dataset aims to fill the gap in multi-track transcription evaluation data for commercial pop music, and to promote the performance enhancement of models in real-world music scenarios.
数据集概述:MulTTiPop
- 全称:MulTTiPop: A Multitrack Transcription Dataset for Pop Music
- 来源:音频片段源自 TheoryTab,对齐的多轨 MIDI 转录源自 Lakh MIDI 数据集。
- 规模:包含 572 个片段,总数据时长约为 3.5 小时。
- 内容:提供对齐的多轨 MIDI 文件及相关元数据文件(含 YouTube 视频 ID 和时间戳)。
- 用途建议:
- 仅获取原始音频的对应片段。
- 仅用于模型评估,不用于训练。
- 数据集划分:按约 3:7 比例分为
dev和test两个子集,不提供训练或验证集。 - 许可证:CC-BY 4.0。
- 下载地址:https://huggingface.co/datasets/gclef-cmu/multtipop
- 相关论文:https://arxiv.org/pdf/2607.08756




