遇见数据集

Saraga: research datasets of Indian Art Music

收藏
Mendeley Data2024-03-27 更新2024-06-28 收录
数据链接:
官方服务:

资源简介:

Dataset introduction This repository contains time aligned melody, rhythm, and structural annotations for two large open corpora of Indian Art Music (Carnatic and Hindustani music). The repository contains Carnatic and Hindustani collections in separated zip files, and each collection is organized by songs grouped by artist concerts/live performances. This organization follows the structure generated by downloading the data using the scripts available at the dataset Github repository: https://github.com/MTG/saraga. Moreover, there is a part of the Carnatic collection, 168 tracks to be specific, that counts with multitrack audio files apart from the mix audio. The considered instruments are: Ghatam, Mridangam, Violin, Voice and Secondary Voice. Annotations in the dataset Section and tempo annotations stored as start and end timestamps together with the name of the section and tempo during the section (in a separate file). Sama annotations referring to rhythmic cycle boundaries stored as timestamps. Phrase annotations stored as timestamps and transcription of the phrases using solfège symbols ({S, r, R, g, G, m, M, P, d, D, n, N}). Audio features automatically extracted and stored: pitch and tonic. For more information about the dataset tracks and annotations, please refer to the Saraga website: https://mtg.github.io/saraga/ Using this dataset We are interested in knowing if you find our datasets useful! If you use our dataset please email us at mtg-info@upf.edu and tell us about your research. *Please note that you can also use this dataset through the MIRDATA library (https://github.com/mir-dataset-loaders/mirdata), where this dataset is in the list of available datasets.

数据集简介 本仓库包含针对两大开源印度古典音乐语料库——卡纳提克音乐(Carnatic)与印度斯坦尼音乐(Hindustani)的时间对齐旋律、节奏与结构标注。本仓库以独立压缩包形式分别存储两类音乐数据集,每个数据集均按艺术家演唱会/现场演出中的曲目进行分组组织。该组织结构遵循通过数据集GitHub仓库(https://github.com/MTG/saraga)提供的脚本下载数据时生成的目录结构。具体而言,卡纳提克音乐数据集包含168条音轨,除混音音频外,还配备有多轨音频文件。本次标注覆盖的乐器包括:加塔姆鼓(Ghatam)、米里丹加鼓(Mridangam)、小提琴(Violin)、人声(Voice)与副人声(Secondary Voice)。 数据集中的标注信息如下:段落与速度标注以起始、结束时间戳形式存储,同时附带段落名称与该段落的速度值,相关数据存储于独立文件中;萨马(Sama)标注用于标记节奏循环边界,以时间戳形式存储;乐句标注以时间戳形式存储,并使用首调唱名符号({S, r, R, g, G, m, M, P, d, D, n, N})对乐句进行转录。 自动提取并存储的音频特征包括:音高与主音。如需了解更多关于数据集音轨与标注的详细信息,请访问Saraga官网:https://mtg.github.io/saraga/ 数据集使用须知:我们十分期待知晓您是否认为本数据集具备使用价值!若您在研究中使用本数据集,请通过邮箱mtg-info@upf.edu联系我们,并告知您的研究内容。*请注意,您也可以通过MIRDATA库(https://github.com/mir-dataset-loaders/mirdata)使用本数据集,该数据集已被纳入该库的可用数据集列表中。

创建时间:
2023-06-28
搜集汇总
数据集介绍
Saraga: research datasets of Indian Art Music 数据集图片
背景与挑战
背景概述
Saraga是一个专注于印度艺术音乐(包括Carnatic和Hindustani风格)的研究数据集,提供时间对齐的旋律、节奏和结构标注,以及自动提取的音频特征如音高和主音。该数据集包含按艺术家音乐会分组的歌曲,其中Carnatic部分还包含168首曲目的多轨音频文件(如Ghatam、Mridangam等乐器),适用于音乐信息检索和分析。数据集发布于2018年,采用开放许可,可通过GitHub和MIRDATA库访问。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务