遇见数据集

FSDKaggle2018

收藏
Zenodo2020-07-28 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

FSDKaggle2018 is an audio dataset containing 11,073 audio files annotated with 41 labels of the AudioSet Ontology. FSDKaggle2018 has been used for the DCASE Challenge 2018 Task 2, which was run as a Kaggle competition titled Freesound General-Purpose Audio Tagging Challenge. <strong>Citation</strong> If you use the FSDKaggle2018 dataset or part of it, please cite our <strong>DCASE 2018 paper</strong>: Eduardo Fonseca, Manoj Plakal, Frederic Font, Daniel P. W. Ellis, Xavier Favory, Jordi Pons, Xavier Serra. "General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline". <em>Proceedings of the DCASE 2018 Workshop</em> (2018) You can also consider citing our <strong>ISMIR 2017 paper</strong>, which describes how we gathered the manual annotations included in FSDKaggle2018. Eduardo Fonseca, Jordi Pons, Xavier Favory, Frederic Font, Dmitry Bogdanov, Andres Ferraro, Sergio Oramas, Alastair Porter, and Xavier Serra, "Freesound Datasets: A Platform for the Creation of Open Audio Datasets", In <em>Proceedings of the 18th International Society for Music Information Retrieval Conference</em>, Suzhou, China, 2017 <strong>Contact</strong> You are welcome to contact Eduardo Fonseca should you have any questions at eduardo.fonseca@upf.edu. <strong>About this dataset</strong> Freesound Dataset Kaggle 2018 (or <strong>FSDKaggle2018</strong> for short) is an audio dataset containing 11,073 audio files annotated with 41 labels of the AudioSet Ontology [1]. FSDKaggle2018 has been used for the Task 2 of the <em>Detection and Classification of Acoustic Scenes and Events</em> (DCASE) Challenge 2018. Please visit the DCASE2018 Challenge Task 2 website for more information. This Task was hosted on the Kaggle platform as a competition titled Freesound General-Purpose Audio Tagging Challenge. It was organized by researchers from the Music Technology Group of Universitat Pompeu Fabra, and from Google Research’s Machine Perception Team. The goal of this competition was to build an audio tagging system that can categorize an audio clip as belonging to one of a set of 41 diverse categories drawn from the AudioSet Ontology. All audio samples in this dataset are gathered from Freesound [2] and are provided here as uncompressed PCM 16 bit, 44.1 kHz, mono audio files. Note that because Freesound content is collaboratively contributed, recording quality and techniques can vary widely. The ground truth data provided in this dataset has been obtained after a data labeling process which is described below in the <em>Data labeling process</em> section. FSDKaggle2018 clips are unequally distributed in the following <strong>41 categories</strong> of the AudioSet Ontology: "Acoustic_guitar", "Applause", "Bark", "Bass_drum", "Burping_or_eructation", "Bus", "Cello", "Chime", "Clarinet", "Computer_keyboard", "Cough", "Cowbell", "Double_bass", "Drawer_open_or_close", "Electric_piano", "Fart", "Finger_snapping", "Fireworks", "Flute", "Glockenspiel", "Gong", "Gunshot_or_gunfire", "Harmonica", "Hi-hat", "Keys_jangling", "Knock", "Laughter", "Meow", "Microwave_oven", "Oboe", "Saxophone", "Scissors", "Shatter", "Snare_drum", "Squeak", "Tambourine", "Tearing", "Telephone", "Trumpet", "Violin_or_fiddle", "Writing". Some other relevant characteristics of FSDKaggle2018: The dataset is split into a train set and a test set. The <strong>train set</strong> is meant to be for system development and includes <strong>~9.5k samples unequally distributed among 41 categories</strong>. The minimum number of audio samples per category in the train set is 94, and the maximum 300. The duration of the audio samples ranges from 300ms to 30s due to the diversity of the sound categories and the preferences of Freesound users when recording sounds. The total duration of the train set is roughly 18h. Out of the ~9.5k samples from the train set, <strong>~3.7k have manually-verified ground truth annotations</strong> and <strong>~5.8k have non-verified annotations</strong>. The non-verified annotations of the train set have a quality estimate of <strong>at least</strong> 65-70% in each category. Checkout the <em>Data labeling process</em> section below for more information about this aspect. Non-verified annotations in the train set are properly flagged in <code>train.csv</code> so that participants can opt to use this information during the development of their systems. The <strong>test set</strong> is composed of <strong>1.6k samples with manually-verified annotations</strong> and with a similar category distribution than that of the train set. The total duration of the test set is roughly 2h. All audio samples in this dataset have a <strong>single label</strong> (i.e. are only annotated with one label). Checkout the <em>Data labeling process </em>section below for more information about this aspect. A single label should be predicted for each file in the test set. <strong>Data labeling process</strong> The data labeling process started from a manual mapping between Freesound tags and AudioSet Ontology categories (or <em>labels</em>), which was carried out by researchers at the Music Technology Group, Universitat Pompeu Fabra, Barcelona. Using this mapping, a number of Freesound audio samples were <strong>automatically annotated</strong> with labels from the AudioSet Ontology. These annotations can be understood as weak labels since they express the presence of a sound category in an audio sample. Then, a <strong>data validation process</strong> was carried out in which a number of participants did listen to the annotated sounds and manually assessed the presence/absence of an automatically assigned sound category, according to the AudioSet category description. Audio samples in FSDKaggle2018 are only annotated with a single ground truth label (see <code>train.csv</code>). A total of <strong>3,710 annotations </strong>included in the train set of FSDKaggle2018 are annotations that have been <strong>manually validated</strong> as present and predominant (some with inter-annotator agreement but not all of them). This means that in most cases there is no additional acoustic material other than the labeled category. In few cases there may be some additional sound events, but these additional events won't belong to any of the 41 categories of FSDKaggle2018. The rest of the annotations have <strong>not</strong> been manually validated and therefore some of them could be inaccurate. Nonetheless, we have <strong>estimated</strong> that <strong>at least</strong> 65-70% of the non-verified annotations per category <strong>in the train set</strong> are indeed correct. It can happen that some of these non-verified audio samples present several sound sources even though only one label is provided as ground truth. These additional sources are typically out of the set of the 41 categories, but in a few cases they could be within. More details about the data labeling process can be found in [3]. <strong>License</strong> FSDKaggle2018 has licenses at two different levels, as explained next. All sounds in Freesound are released under Creative Commons (CC) licenses, and each audio clip has its own license as defined by the audio clip uploader in Freesound. For attribution purposes and to facilitate attribution of these files to third parties, we include a relation of the audio clips included in FSDKaggle2018 and their corresponding license. The licenses are specified in the files <code>train_post_competition.csv</code> and <code>test_post_competition_scoring_clips.csv</code>. In addition, FSDKaggle2018 as a whole is the result of a curation process and it has an additional license. FSDKaggle2018 is released under CC-BY. This license is specified in the <code>LICENSE-DATASET</code> file downloaded with the <code>FSDKaggle2018.doc</code> zip file. <strong>Files</strong> FSDKaggle2018 can be downloaded as a series of zip files with the following directory structure: <pre>root │ └───FSDKaggle2018.audio_train/ Audio clips in the train set │ └───FSDKaggle2018.audio_test/ Audio clips in the test set │ └───FSDKaggle2018.meta/ Files for evaluation setup │ │ │ └───train_post_competition.csv Data split and ground truth for the train set │ │ │ └───test_post_competition_scoring_clips.csv Ground truth for the test set │ └───FSDKaggle2018.doc/ │ └───README.md The dataset description file you are reading │ └───LICENSE-DATASET License of FSDKaggle2018 dataset as a whole </pre> <strong>NOTE</strong>: the original <code>train.csv</code> file provided during the competition has been updated with more metadata (licenses, Freesound ids, etc.) into <code>train_post_competition.csv</code>. Likewise, the original <code>test.csv</code> that was not public during the competition is now available with ground truth and metadata as <code>test_post_competition_scoring_clips.csv</code>. The file name <code>test_post_competition_scoring_clips.csv</code> refers to the fact that only the 1600 clips used for systems' ranking are included. During the competition, an additional subset of <em>padding</em> clips was added in order to prevent undesired practices. This <em>padding</em> subset (that was never used for systems' ranking) is no longer included in the dataset (see our DCASE 2018 paper for more details.) Each row (i.e. audio clip) of the <code>train_post_competition.csv</code> file contains the following information: <code>fname</code>: the file name <code>label</code>: the audio classification label (ground truth) <code>manually_verified</code>: Boolean (1 or 0) flag to indicate whether or not that annotation has been manually verified; see description above for more info <code>freesound_id</code>: the Freesound id for the audio clip <code>license</code>: the license for the audio clip Each row (i.e. audio clip) of the <code>test_post_competition_scoring_clips.csv</code> file contains the following information: <code>fname</code>: the file name <code>label</code>: the audio classification label (ground truth) <code>usage</code>: string that indicates to which Kaggle leaderboard the clip was associated during the competition: <code>Public</code> or <code>Private</code> <code>freesound_id</code>: the Freesound id for the audio clip <code>license</code>: the license for the audio clip <strong>Baseline System</strong> A CNN baseline system for FSDKaggle2018 is available at https://github.com/DCASE-REPO/dcase2018_baseline/tree/master/task2. <strong>References and links</strong> [1] Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. "Audio set: An ontology and human-labeled dartaset for audio events." Proceedings of the Acoustics, Speech and Signal Processing International Conference, 2017. [2] Frederic Font, Gerard Roma, and Xavier Serra. "Freesound technical demo." Proceedings of the 21st ACM international conference on Multimedia, 2013. https://freesound.org [3] Eduardo Fonseca, Jordi Pons, Xavier Favory, Frederic Font, Dmitry Bogdanov, Andres Ferraro, Sergio Oramas, Alastair Porter, and Xavier Serra. "Freesound Datasets: A Platform for the Creation of Open Audio Datasets." Proceedings of the International Conference on Music Information Retrieval, 2017. PDF here Freesound Annotator: https://annotator.freesound.org/<br> Freesound: https://freesound.org<br> Eduardo Fonseca's personal website: http://www.eduardofonseca.net/<br> More datasets collected by us: http://www.eduardofonseca.net/datasets/ <strong>Acknowledgments</strong> This work is partially supported by the European Union’s Horizon 2020 research and innovation programme under grant agreement No 688382 AudioCommons. Eduardo Fonseca is also sponsored by a Google Faculty Research Award 2017. We thank everyone who contributed to FSDKaggle2018 with annotations.

FSDKaggle2018是一款音频数据集,包含11073个音频文件,标注有来自AudioSet本体(AudioSet Ontology)的41个类别标签。该数据集曾被用于2018年声学场景与事件检测与分类挑战赛(DCASE Challenge 2018)任务2,该任务以名为“Freesound通用音频标签挑战赛”的Kaggle竞赛形式举办。 **引用** 若您使用FSDKaggle2018数据集或其部分内容,请引用我们的**DCASE 2018论文**: Eduardo Fonseca, Manoj Plakal, Frederic Font, Daniel P. W. Ellis, Xavier Favory, Jordi Pons, Xavier Serra. "General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline". *Proceedings of the DCASE 2018 Workshop* (2018) 您也可以引用我们的**ISMIR 2017论文**,该论文详述了FSDKaggle2018数据集人工标注的采集流程: Eduardo Fonseca, Jordi Pons, Xavier Favory, Frederic Font, Dmitry Bogdanov, Andres Ferraro, Sergio Oramas, Alastair Porter, and Xavier Serra, "Freesound Datasets: A Platform for the Creation of Open Audio Datasets", In *Proceedings of the 18th International Society for Music Information Retrieval Conference*, Suzhou, China, 2017 **联系我们** 若您有任何疑问,欢迎联系Eduardo Fonseca,邮箱:eduardo.fonseca@upf.edu。 **数据集概况** Freesound Kaggle 2018数据集(简称FSDKaggle2018)是一款音频数据集,包含11073个音频文件,标注有来自AudioSet本体(AudioSet Ontology)的41个类别标签[1]。FSDKaggle2018曾被用于2018年声学场景与事件检测与分类挑战赛(DCASE Challenge 2018)的任务2,更多信息可访问DCASE2018挑战赛任务2官网。该任务以Kaggle平台上的“Freesound通用音频标签挑战赛”竞赛形式举办,由庞培法布拉大学音乐技术组与谷歌研究机器学习感知团队的研究者共同组织。 本次竞赛的目标是构建音频标注系统,可将音频片段归类为从AudioSet本体中选取的41个多样化类别之一。 本数据集的所有音频样本均采集自Freesound平台(Freesound),并以未压缩的PCM 16位、44.1kHz单声道音频文件形式提供。请注意,由于Freesound平台的内容由用户协作上传,录音质量与录制手法存在较大差异。 本数据集提供的真实标签数据经过了下述*数据标注流程*章节中详述的标注流程处理。FSDKaggle2018的音频片段在AudioSet本体的以下**41个类别**中分布不均: "Acoustic_guitar", "Applause", "Bark", "Bass_drum", "Burping_or_eructation", "Bus", "Cello", "Chime", "Clarinet", "Computer_keyboard", "Cough", "Cowbell", "Double_bass", "Drawer_open_or_close", "Electric_piano", "Fart", "Finger_snapping", "Fireworks", "Flute", "Glockenspiel", "Gong", "Gunshot_or_gunfire", "Harmonica", "Hi-hat", "Keys_jangling", "Knock", "Laughter", "Meow", "Microwave_oven", "Oboe", "Saxophone", "Scissors", "Shatter", "Snare_drum", "Squeak", "Tambourine", "Tearing", "Telephone", "Trumpet", "Violin_or_fiddle", "Writing". FSDKaggle2018的其他相关特性如下: 1. 数据集划分为训练集与测试集。 2. **训练集**用于模型开发,包含**约9500个分布不均的41个类别的音频样本**。训练集每个类别的音频样本最少为94个,最多为300个。由于声音类别的多样性以及Freesound平台用户的录制偏好,音频样本的时长范围为300ms至30s。训练集总时长约为18小时。 3. 在训练集的约9500个样本中,**约3700个带有经人工验证的真实标注**,**约5800个带有未验证标注**。训练集的未验证标注在每个类别中的准确率估计至少为65%-70%。有关该部分的更多信息,请参阅下述*数据标注流程*章节。 4. 训练集的未验证标注会在<code>train.csv</code>中进行明确标记,以便参赛者在开发模型时选择是否使用该类标注信息。 5. **测试集**包含**1600个带有经人工验证标注的样本**,其类别分布与训练集相似。测试集总时长约为2小时。 6. 本数据集的所有音频样本均带有**单个标签**(即仅用一个标签进行标注)。有关该部分的更多信息,请参阅下述*数据标注流程*章节。测试集的每个文件仅需预测单个标签。 **数据标注流程** 数据标注流程始于巴塞罗那庞培法布拉大学音乐技术组的研究者完成的Freesound标签与AudioSet本体类别(或称**标签**)之间的手动映射。基于该映射,大量Freesound音频样本被**自动标注**上来自AudioSet本体的标签。这类标注可被视为弱标注,因为它们仅表明音频片段中存在某类声音。 随后,研究者开展了**数据验证流程**:让多名参与者聆听已标注的音频片段,并根据AudioSet类别的描述手动评估自动分配的声音类别是否存在。 FSDKaggle2018中的音频片段仅带有单个真实标签(详见<code>train.csv</code>)。训练集中共有**3710个标注**经过**人工验证**,确认对应声音存在且为主要声源(部分标注存在标注者间一致性,部分则没有)。这意味着在大多数情况下,音频片段中除标注类别外无其他声学内容;少数情况下可能存在额外的声音事件,但这些额外事件不属于FSDKaggle2018的41个类别之一。 其余标注均未经过人工验证,因此其中部分标注可能不准确。尽管如此,我们估计训练集每个类别中**至少65%-70%的未验证标注**确实正确。部分未验证的音频片段可能包含多个声源,尽管仅提供了一个标签作为真实标注。这些额外声源通常不属于41个类别之列,但少数情况下也可能属于其中。 有关数据标注流程的更多细节可参阅文献[3]。 **使用许可** FSDKaggle2018分为两个层级的许可,如下所述。 Freesound平台上的所有音频均采用知识共享(Creative Commons,简称CC)许可,每个音频片段的许可由其上传者在Freesound平台上定义。为便于归因并方便将这些文件的归因告知第三方,我们在FSDKaggle2018数据集中附带了音频片段与其对应许可的对应关系。相关许可信息存储在<code>train_post_competition.csv</code>和<code>test_post_competition_scoring_clips.csv</code>文件中。 此外,FSDKaggle2018作为整体经过了整理流程,带有额外的使用许可。FSDKaggle2018整体采用CC-BY许可,相关许可信息存储在与<code>FSDKaggle2018.doc</code>压缩包一同下载的<code>LICENSE-DATASET</code>文件中。 **文件结构** FSDKaggle2018可通过一系列压缩包下载,目录结构如下: <pre> root │ └───FSDKaggle2018.audio_train/ 训练集音频片段 │ └───FSDKaggle2018.audio_test/ 测试集音频片段 │ └───FSDKaggle2018.meta/ 评估配置文件 │ │ │ └───train_post_competition.csv 训练集数据拆分与真实标签文件 │ │ │ └───test_post_competition_scoring_clips.csv 测试集真实标签文件 │ └───FSDKaggle2018.doc/ │ └───README.md 本数据集说明文件(即您正在阅读的文档) │ └───LICENSE-DATASET FSDKaggle2018数据集整体许可协议 </pre> **注意事项**:竞赛期间提供的原始<code>train.csv</code>文件已更新,新增了更多元数据(许可信息、Freesound ID等),更新后的文件为<code>train_post_competition.csv</code>。同样,竞赛期间未公开的原始<code>test.csv</code>文件现已更名为<code>test_post_competition_scoring_clips.csv</code>,并附带了真实标签与元数据。该文件名之所以为<code>test_post_competition_scoring_clips.csv</code>,是因为该文件仅包含用于模型排名的1600个音频片段。竞赛期间还添加了额外的**填充**片段子集,以防止不当操作,但该填充子集(从未用于模型排名)现已从数据集中移除(更多细节可参阅我们的DCASE 2018论文)。 <code>train_post_competition.csv</code>文件的每一行(对应一个音频片段)包含以下信息: - <code>fname</code>:文件名 - <code>label</code>:音频分类标签(真实标签) - <code>manually_verified</code>:布尔值(1或0),用于标记该标注是否经过人工验证;详见前文描述 - <code>freesound_id</code>:该音频片段的Freesound平台ID - <code>license</code>:该音频片段的使用许可 <code>test_post_competition_scoring_clips.csv</code>文件的每一行(对应一个音频片段)包含以下信息: - <code>fname</code>:文件名 - <code>label</code>:音频分类标签(真实标签) - <code>usage</code>:字符串,用于标记该音频片段在竞赛期间对应的Kaggle排行榜类型:<code>Public</code>(公共榜)或<code>Private</code>(私有榜) - <code>freesound_id</code>:该音频片段的Freesound平台ID - <code>license</code>:该音频片段的使用许可 **基线系统** 针对FSDKaggle2018的卷积神经网络(Convolutional Neural Network,简称CNN)基线系统可在以下地址获取:https://github.com/DCASE-REPO/dcase2018_baseline/tree/master/task2。 **参考文献与链接** [1] Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. "Audio set: An ontology and human-labeled dataset for audio events." Proceedings of the Acoustics, Speech and Signal Processing International Conference, 2017. [2] Frederic Font, Gerard Roma, and Xavier Serra. "Freesound technical demo." Proceedings of the 21st ACM international conference on Multimedia, 2013. https://freesound.org [3] Eduardo Fonseca, Jordi Pons, Xavier Favory, Frederic Font, Dmitry Bogdanov, Andres Ferraro, Sergio Oramas, Alastair Porter, and Xavier Serra. "Freesound Datasets: A Platform for the Creation of Open Audio Datasets." Proceedings of the International Conference on Music Information Retrieval, 2017. 可在此获取PDF文件 Freesound标注工具:https://annotator.freesound.org/<br> Freesound平台:https://freesound.org<br> Eduardo Fonseca个人主页:http://www.eduardofonseca.net/<br> 我们收集的更多数据集:http://www.eduardofonseca.net/datasets/ **致谢** 本工作部分受到欧盟地平线2020研究与创新计划的资助,项目编号为688382 AudioCommons。Eduardo Fonseca同时获得了2017年谷歌教师研究奖的资助。我们感谢所有为FSDKaggle2018提供标注的参与者。

提供机构:
Zenodo
创建时间:
2019-01-30
二维码
社区交流群
二维码
科研交流群
商业服务