Covers and Hummings Aligned Dataset (CHAD)
收藏资源简介:
Covers and Hummings Aligned Dataset (CHAD)是由华为诺亚方舟实验室创建的一个新型数据集,专注于音乐信息检索中的Query-by-Humming任务。该数据集包含5494首原曲和31630首翻唱版本,以及5164个哼唱片段,总计81781个音频片段,时长超过270小时。CHAD数据集通过精确的时间对齐技术,确保哼唱片段与原曲版本在时间结构上的一致性。创建过程中,利用了众包和半监督学习方法,有效地收集和扩展了数据集。该数据集主要应用于音乐检索系统,旨在通过用户哼唱的旋律快速准确地找到对应的歌曲,解决了传统音乐搜索系统中用户需精确记忆歌词或播放完整歌曲的问题。
The Covers and Hummings Aligned Dataset (CHAD) is a novel dataset developed by Huawei Noah's Ark Lab, targeting the Query-by-Humming (QbH) task in the field of Music Information Retrieval (MIR). It consists of 5,494 original songs, 31,630 cover versions, and 5,164 humming clips, with a total of 81,781 audio segments and an overall duration exceeding 270 hours. The CHAD dataset employs precise time alignment technologies to guarantee the temporal structural consistency between each humming clip and its corresponding original song version. During its curation, crowdsourcing and semi-supervised learning approaches were adopted to efficiently collect and expand the dataset. This dataset is primarily utilized in music retrieval systems, with the goal of rapidly and accurately locating matching songs based on user-hummed melodies, thereby addressing the limitation of traditional music search systems that demand users to precisely recall lyrics or play full songs to conduct a search.

- 1A Semi-Supervised Deep Learning Approach to Dataset Collection for Query-By-Humming Task华为诺亚方舟实验室 · 2023年



