遇见数据集

SargamBench

收藏
Zenodo2026-08-12 更新2026-08-13 收录
官方服务:

资源简介:

SargamBench is a dataset of lyric-conditioned Sargam sequences for Hindustani classical music. It contains 1,026 validated lyric–Sargam pairs, retained from 1,037 candidates screened across compositions in Hindi, Sanskrit, and Bengali. The released transcriptions, normalised symbolic representations, and syllable-level lyric–melody alignments were prepared by the research team from publicly accessible educational and archival notation sources. Each entry pairs a syllable-level lyric sequence with a melody sequence represented using a standardised 33-token Sargam vocabulary and is linked to the raga profile of the corresponding composition for rule-based symbolic validation. The dataset covers ten selected raga profiles, with one profile associated with each of the ten thaats in the Bhatkhande classification system. Along with the core pairs, the release includes machine-readable raga-rule files, composition-level train, validation, and test split assignments, binary sequence-validity labels, and evaluation scripts for two benchmark tasks: end-sequence completion and melody infilling. All entries include both Devanagari and Latin representations to support researchers unfamiliar with Indic scripts. No audio, MIDI, or tala data are included in this release. Most existing Indian music datasets are audio-oriented and do not provide token-level lyric–melody alignment together with raga-validity labels. SargamBench addresses this gap and is intended to support research on music language models, symbolic melody generation, raga-constrained sequence evaluation, and computational ethnomusicology.

提供机构:
Zenodo
创建时间:
2026-08-12
二维码
社区交流群
二维码
科研交流群
商业服务