遇见数据集

Spotify Million Playlist: Recsys Challenge 2018 Dataset

收藏
NIAID Data Ecosystem2026-03-13 收录
数据链接:
官方服务:

资源简介:

Spotify Million Playlist Dataset Challenge Summary The Spotify Million Playlist Dataset Challenge consists of a dataset and evaluation to enable research in music recommendations. It is a continuation of the RecSys Challenge 2018, which ran from January to July 2018. The dataset contains 1,000,000 playlists, including playlist titles and track titles, created by users on the Spotify platform between January 2010 and October 2017. The evaluation task is automatic playlist continuation: given a seed playlist title and/or initial set of tracks in a playlist, to predict the subsequent tracks in that playlist. This is an open-ended challenge intended to encourage research in music recommendations, and no prizes will be awarded (other than bragging rights). Background Playlists like Today’s Top Hits and RapCaviar have millions of loyal followers, while Discover Weekly and Daily Mix are just a couple of our personalized playlists made especially to match your unique musical tastes. Our users love playlists too. In fact, the Digital Music Alliance, in their 2018 Annual Music Report, state that 54% of consumers say that playlists are replacing albums in their listening habits. But our users don’t love just listening to playlists, they also love creating them. To date, over 4 billion playlists have been created and shared by Spotify users. People create playlists for all sorts of reasons: some playlists group together music categorically (e.g., by genre, artist, year, or city), by mood, theme, or occasion (e.g., romantic, sad, holiday), or for a particular purpose (e.g., focus, workout). Some playlists are even made to land a dream job, or to send a message to someone special. The other thing we love here at Spotify is playlist research. By learning from the playlists that people create, we can learn all sorts of things about the deep relationship between people and music. Why do certain songs go together? What is the difference between “Beach Vibes” and “Forest Vibes”? And what words do people use to describe which playlists? By learning more about nature of playlists, we may also be able to suggest other tracks that a listener would enjoy in the context of a given playlist. This can make playlist creation easier, and ultimately help people find more of the music they love. Dataset To enable this type of research at scale, in 2018 we sponsored the RecSys Challenge 2018, which introduced the Million Playlist Dataset (MPD) to the research community. Sampled from the over 4 billion public playlists on Spotify, this dataset of 1 million playlists consist of over 2 million unique tracks by nearly 300,000 artists, and represents the largest public dataset of music playlists in the world. The dataset includes public playlists created by US Spotify users between January 2010 and November 2017. The challenge ran from January to July 2018, and received 1,467 submissions from 410 teams. A summary of the challenge and the top scoring submissions was published in the ACM Transactions on Intelligent Systems and Technology. In September 2020, we re-released the dataset as an open-ended challenge on AIcrowd.com. The dataset can now be downloaded by registered participants from the Resources page. Each playlist in the MPD contains a playlist title, the track list (including track IDs and metadata), and other metadata fields (last edit time, number of playlist edits, and more). All data is anonymized to protect user privacy. Playlists are sampled with some randomization, are manually filtered for playlist quality and to remove offensive content, and have some dithering and fictitious tracks added to them. As such, the dataset is not representative of the true distribution of playlists on the Spotify platform, and must not be interpreted as such in any research or analysis performed on the dataset. Dataset Contains 1000 examples of each scenario: Title only (no tracks) Title and first track Title and first 5 tracks First 5 tracks only Title and first 10 tracks First 10 tracks only Title and first 25 tracks Title and 25 random tracks Title and first 100 tracks Title and 100 random tracks Download Link Full Details: https://www.aicrowd.com/challenges/spotify-million-playlist-dataset-challenge Download Link: https://www.aicrowd.com/challenges/spotify-million-playlist-dataset-challenge/dataset_files

Spotify百万播放列表数据集挑战赛 摘要 Spotify百万播放列表数据集挑战赛包含一套数据集与评估框架,旨在推动音乐推荐领域的研究。该赛事是2018年RecSys挑战赛的延续,原赛事于2018年1月至7月举办。数据集包含100万条用户于2010年1月至2017年10月期间在Spotify平台上创建的播放列表,涵盖播放列表标题与曲目标题信息。本次评估任务为自动播放列表续接:给定初始播放列表标题和/或部分初始曲目,预测该播放列表后续应添加的曲目。本挑战赛为开放式研究挑战,旨在鼓励音乐推荐领域的研究探索,除荣誉外不设置任何奖励。 背景 诸如《今日热门金曲》(Today’s Top Hits)与《说唱Caviar》(RapCaviar)这类播放列表拥有数百万忠实听众,而《发现周报》(Discover Weekly)与《每日精选混音》(Daily Mix)则是专为匹配用户独特音乐品味打造的个性化播放列表。 用户对播放列表的喜爱不仅限于收听,更热衷于创作。事实上,数字音乐联盟(Digital Music Alliance)在2018年度音乐报告中指出,54%的消费者表示播放列表正在取代专辑成为他们的主流收听载体。 截至目前,Spotify用户已创建并分享超过40亿条播放列表。人们创作播放列表的动机多种多样:部分播放列表按类别归类音乐(如按流派、艺人、发行年份或城市划分),或基于情绪、主题、场合(如浪漫、伤感、节日)分类,亦或是为特定场景打造(如专注、健身)。甚至有播放列表被用于助力求职,或是向心仪之人传递心意。 Spotify同样重视播放列表相关研究。通过分析用户创建的播放列表,我们能够深入探索人类与音乐之间的深层关联:为何某些曲目会天然适配?“海滩氛围”与“森林氛围”两类播放列表的差异何在?人们又会用哪些词汇来定义不同的播放列表? 深入了解播放列表的本质特征,还有助于为特定播放列表推荐听众可能喜爱的新增曲目,这不仅能简化播放列表的创作流程,最终也能帮助用户发现更多心仪的音乐。 数据集 为支持大规模相关研究,Spotify于2018年发起2018年RecSys挑战赛,并向研究社区推出百万播放列表数据集(Million Playlist Dataset,MPD)。该数据集从Spotify平台上超过40亿条公开播放列表中采样而来,包含100万条播放列表,涵盖近30万位艺人的超200万首独特曲目,是目前全球规模最大的公开音乐播放列表数据集。数据集收录的播放列表均为2010年1月至2017年11月期间由美国Spotify用户创建的公开播放列表。2018年的挑战赛于1月至7月举办,共收到来自410支团队的1467份参赛作品,挑战赛总结与顶尖参赛方案已发表于《ACM智能系统与技术汇刊》(ACM Transactions on Intelligent Systems and Technology)。 2020年9月,Spotify将该数据集重新以开放式挑战赛的形式发布于AIcrowd.com平台。已注册的参赛者可通过资源页面下载该数据集。 百万播放列表数据集中的每条播放列表均包含播放列表标题、曲目列表(含曲目ID与元数据)以及其他元数据字段(如最后编辑时间、播放列表编辑次数等)。所有数据均已做匿名化处理以保护用户隐私。播放列表采样过程引入了随机化机制,同时会人工筛选以保证播放列表质量并移除违规内容,还会添加部分抖动数据与虚构曲目。因此,该数据集无法代表Spotify平台上播放列表的真实分布,任何基于该数据集开展的研究与分析均不应将其作为真实分布的参照。 数据集涵盖场景 本次挑战赛提供以下10类测试场景各1000个示例: 1. 仅提供播放列表标题(无曲目) 2. 提供播放列表标题与第一首曲目 3. 提供播放列表标题与前五首曲目 4. 仅提供前五首曲目 5. 提供播放列表标题与前十首曲目 6. 仅提供前十首曲目 7. 提供播放列表标题与前二十五首曲目 8. 提供播放列表标题与二十五首随机曲目 9. 提供播放列表标题与前一百首曲目 10. 提供播放列表标题与一百首随机曲目 下载链接 完整详情:https://www.aicrowd.com/challenges/spotify-million-playlist-dataset-challenge 数据集下载:https://www.aicrowd.com/challenges/spotify-million-playlist-dataset-challenge/dataset_files

创建时间:
2022-04-09
二维码
社区交流群
二维码
科研交流群
商业服务