Songs
收藏arXiv2025-09-30 收录
官方服务:
资源简介:
该数据集名为SG,其中包含了歌曲实体,部分实体指向同一首歌曲。实验中,我们对同一表格中的歌曲条目进行了匹配。在匹配过程中,我们使用了以下相似度度量方法:对于歌曲标题,采用了Jaccard相似度;对于歌曲标题和发行信息,则使用了Jaro-Winkler距离;而对于时长这一属性,则采用了数字相似度。该数据集的规模为8,312对歌曲实体,其中1,412对被认为是等效的。所执行的任务是实体解析。
This dataset is named SG, which contains song entities, and some entities refer to the same song. In the experiment, we performed matching on song entries within the same table. During the matching process, the following similarity metrics were adopted: Jaccard similarity for song titles; Jaro-Winkler distance for song titles and release information; and numerical similarity for the duration attribute. The dataset consists of 8,312 pairs of song entities, of which 1,412 pairs are considered equivalent. The task executed is entity resolution.
搜集汇总
数据集介绍

背景与挑战
背景概述
该数据集名为'Songs',包含两个CSV文件:matches_msd_msd.csv(17M)和msd.csv(100M),可能涉及歌曲信息或匹配数据,但详情页未提供具体描述,因此无法进一步总结其内容或特点。
以上内容由遇见数据集搜集并总结生成



