遇见数据集

DiVers: A large and diverse version identification dataset.

收藏
Zenodo2025-12-22 更新2026-05-26 收录
官方服务:

资源简介:

This version identification dataset is based on Discogs-VI and contains more than 700k additional versions from YouTube. In this repository, we provide three datasets with compatible splits: dvi2: a second, cleaned version of Discogs-VI yvi: a YouTube crawl comprising newly found versions on YouTube divers1m: the combination of both datasets File Structure we include a subdir with Torch and one with JSON files with the same data each dataset is a dictionary with the following content: info: global version IDs mapping to the version metadata global ID is the clique ID concatenated with the version ID with a ":" in between. Example: C-01:V-07 with clique ID C-01 and version ID V0-7 depending on the type of file (indicated by the sub-directory), we either have light: only basic metadata on the song level and the sampling rate and length rich: everything in light but also the matched concepts, YouTube title and tags, etc. split: split key (train, valid, test) mapping to the clique identifiers which map to lists of version identifiers

本版本识别数据集基于Discogs-VI构建,额外收录了来自YouTube的70余万条曲目版本。本仓库共提供三个兼容统一划分规则的数据集: 1. dvi2:Discogs-VI的第二代净化版本 2. yvi:从YouTube爬取得到的全新曲目版本数据集 3. divers1m:上述两个数据集的合并版本 ### 文件结构 本仓库包含两个子目录:其一存储适配PyTorch的数据集文件,其二存储与前者数据完全一致的JSON格式数据集文件。 每个数据集均为字典结构,包含以下字段: - `info`:全局版本ID映射至版本元数据。其中全局ID由组ID与版本ID以冒号拼接而成,示例格式为`C-01:V-07`,其中`C-01`为组ID,`V-07`为版本ID。 根据文件所在子目录的不同,数据集分为两类: 1. 轻量型(light):仅包含歌曲级基础元数据、采样率与音频时长 2. 丰富型(rich):包含轻量型的全部字段,同时额外收录匹配到的概念、YouTube视频标题与标签等信息 - `split`:划分键(对应训练集、验证集、测试集)映射至组标识符,组标识符进一步映射至版本标识符列表

提供机构:
Zenodo
创建时间:
2025-08-29
二维码
社区交流群
二维码
科研交流群
商业服务