MusiCorpus
收藏资源简介:
MusiCorpus是由计算机视觉中心、巴塞罗那自治大学及摩拉维亚图书馆等机构联合创建的大型历史手写乐谱数据集,旨在解决光学音乐识别领域缺乏真实训练数据的关键瓶颈。该数据集包含1,309页源自欧洲档案机构的原始乐谱扫描图像,涵盖管弦乐谱、分谱及钢琴谱等多种类型,并提供了MusicXML转录文本与符号级标注,数据总量达数十万音乐符号。其创建过程通过与多个文化遗产机构合作,采用专家手工标注与标准化编码流程,确保了数据的多样性与学术严谨性。本数据集主要应用于训练和评估端到端及基于目标检测的光学音乐识别系统,以推动历史音乐文献的自动化转录与数字化保存,为音乐学研究和文化遗产保护提供关键技术支撑。
MusiCorpus is a large-scale historical handwritten musical score dataset jointly created by institutions including the Computer Vision Center, Autonomous University of Barcelona, and Moravian Library. It aims to address the critical bottleneck of lacking real-world training data in the field of optical music recognition (OMR). The dataset contains 1,309 pages of original scanned musical score images sourced from European archival institutions, covering various types such as orchestral scores, individual performance parts, and piano scores. It also provides MusicXML transcriptions and symbol-level annotations, with a total of hundreds of thousands of musical symbols in the corpus. During its development, the dataset was collaboratively constructed with multiple cultural heritage institutions, adopting expert manual annotation and standardized encoding procedures to ensure data diversity and academic rigor. This dataset is mainly used for training and evaluating end-to-end and object detection-based optical music recognition systems, so as to promote automated transcription and digital preservation of historical musical documents, and provide critical technical support for musicological research and cultural heritage conservation.

- 1A Dataset for the Recognition of Historical and Handwritten Music Scores in Western Notation计算机视觉中心; 巴塞罗那自治大学·计算机科学系; 巴塞罗那自治大学·艺术与音乐学系; 查尔斯大学·形式与应用语言学研究所; 摩拉维亚图书馆 · 2026年



