Raga Ornamentation Detection (ROD)
收藏资源简介:
The Raga Ornamentation Detection (ROD) dataset is composed of 212 audio files of varying lengths recorded by two expert musicians. Singer 1 performed 108 recordings, while singer 2 produced 104. In total, the data set contains 4.08 hours of audio. In addition to the audio files, the dataset also includes accompaniments like drone Tanpura and percussion (Tabla). The audio files are labelled in the strong labelling fashion. For more details, check Readme.mdEmail: krsumit@iitk.ac.in for further queries ROD Dataset: Ornamentation Detection in Indian Classical Vocal Music ROD Dataset: Ornamentation Detection in Indian Classical Vocal Music Authors: Sumit Kumar, et al.Affiliation: Indian Institute of Technology KanpurCorresponding Paper: Recognizing Ornamentations in Singing Voices in Indian Art MusicLink to paper: https://ieeexplore.ieee.org/document/11275638License: Creative Commons Attribution 4.0 International (CC BY 4.0) Overview: The Raga Ornamentation Detection (ROD) dataset accompanies the paper titled “Recognizing Ornamentations in Singing Voices in Indian Art Music”. It comprises professionally recorded and annotated vocal performances collected from two distinct sources: Expert Teachers: Studio-quality recordings of trained Indian classical vocalists. Prasar Bharati Broadcast Recordings: Real-world concert and archival recordings. Each recording is annotated for prominent ornamentations used in Hindustani vocal music, such as meend, kan, andolan, murki, gamak, and nyas swar.For reproducibility and benchmarking of machine learning models, the dataset includes raw audio recordings as well as precomputed spectrogram features. Data Description: Audio Recordings Sampling rate: 44.1 kHz Bit depth: 16-bit PCM Channels: Mono Format: WAV Duration: 2–10 minutes per recording Sources: Teacher lessons and Prasar Bharati broadcast renditions Spectrograms Extraction library: [Parselmouth (Praat Python Interface)](https://parselmouth.readthedocs.io/) Sampling rate (Fs): 16,000 Hz Window length (win_length): 0.035 sec FFT window size (n_fft): 4096 Hop length (hop_length): 0.0175 sec Maximum frequency (f_max): 5000 Hz Clip duration: 10 sec Number of classes: 7 Spectral resolution: 120 frequency bins (n_chroma = 120) Annotations Two levels of annotations are provided in the dataset. Event-level annotations are stored in plain text (.txt) format. Each annotation file corresponds to a single audio clip and contains the onset time, offset time, and ornamentation label, separated by tab characters. <onset_time> <offset_time> <ornament_label> Onset (s) Offset (s) Label 12.186986 12.632982 K 12.785793 13.857179 H 13.964114 14.629196 H 15.093979 15.925650 Me 17.727695 18.842940 Mu 19.276945 23.220490 H Notations: | Notation | Ornament | | K / k | Kanswar | | M / Me / Me1 | Meend | | An / AN / An1 | Andolan | | Mu / MU / Mu1 | Murki | | H / h / H1| Holding Note | | G / g | Gamak | Potential Applications: - music pedagogy - expressive singing voice synthesis - singer identification - content-based music recommendation - source separation - tonic identification - tala identification - raga identification Citation: If you use this dataset, please cite the following: S. Kumar, P. Singh and V. Arora, "Recognizing Ornaments in Vocal Indian Art Music With Active Annotation," in IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pp. 220-229, 2026, doi: 10.1109/TASLPRO.2025.3639738. BibTex: @ARTICLE{11275638, author={Kumar, Sumit and Singh, Parampreet and Arora, Vipul}, journal={IEEE Transactions on Audio, Speech and Language Processing}, title={Recognizing Ornaments in Vocal Indian Art Music With Active Annotation}, year={2026}, volume={34}, number={}, pages={220-229}, doi={10.1109/TASLPRO.2025.3639738}} License: This dataset is released under the **Creative Commons Attribution 4.0 International License (CC BY 4.0)**. You are free to **use, share, and modify** the data, provided that appropriate **credit is given** to the authors and the original source. Related Links : 📄 [Paper on IEEExplore] 📄 [Paper on arXiv] 🎵 [Dataset] 💻 [GitHub Repository]



