MACS - Multi-Annotator Captioned Soundscapes
收藏资源简介:
This is a dataset containing <strong>audio captions </strong>and corresponding<strong> audio tags</strong> for a number of 3930 audio files of the TAU Urban Acoustic Scenes 2019 development dataset (airport, public square, and park). The files were annotated using a web-based tool. Each file is annotated by multiple annotators that provided tags and a one-sentence description of the audio content. The data also includes<strong> annotator competence </strong>estimated using MACE (Multi-Annotator Competence Estimation). The annotation procedure, processing and analysis of the data are presented in the following papers: Irene Martin-Morato, Annamaria Mesaros. <em>What is the ground truth? Reliability of multi-annotator data for audio tagging, </em>29th European Signal Processing Conference, EUSIPCO 2021 Irene Martin-Morato, Annamaria Mesaros. <em>Diversity and bias in audio captioning datasets, </em>submitted to DCASE 2021 Workshop (to be updated with arxiv link) Data is provided as two files: <strong>MACS.yaml</strong> - containing the complete annotations in the following format: - filename: file1.wav <br> annotations:<br> - annotator_id: ann_1<br> sentence: caption text<br> tags:<br> - tag1<br> - tag2 <br> - annotator_id: ann_2<br> sentence: caption text<br> tags:<br> - tag1 <strong>MACS_competence.csv</strong> - containing the estimated annotator competence; for each annotator_id in the yaml file, competence is a number between 0 (considered as annotating at random) and 1 id [tab] competence The audio files can be downloaded from https://zenodo.org/record/2589280 and are covered by their own license.
本数据集包含TAU城市声学场景2019开发数据集(TAU Urban Acoustic Scenes 2019 Development Dataset)中3930个音频文件的音频字幕(audio captions)与对应音频标签(audio tags),涵盖机场、公共广场与公园三类场景。所有音频文件均通过基于网页的标注工具完成标注,每个文件由多名标注者进行标注,标注内容包含音频标签与一段描述音频内容的单句说明。数据集还包含使用MACE(Multi-Annotator Competence Estimation,多标注者能力评估)方法估算得到的标注者能力(annotator competence)数据。 本数据集的标注流程、数据处理与分析方法详见以下两篇论文: 1. Irene Martin-Morato、Annamaria Mesaros. 《何为真实标签?音频标注多标注者数据的可靠性》,发表于第29届欧洲信号处理会议(EUSIPCO 2021) 2. Irene Martin-Morato、Annamaria Mesaros. 《音频字幕数据集的多样性与偏差》,已提交至DCASE 2021研讨会(后续将补充arXiv链接) 数据集以两个文件形式提供: - MACS.yaml:包含完整标注信息,格式如下: - 文件名:file1.wav 标注项: - 标注者ID:ann_1 句子:字幕文本 标签: - tag1 - tag2 - 标注者ID:ann_2 句子:字幕文本 标签: - tag1 - MACS_competence.csv:包含估算得到的标注者能力数据;对于YAML文件中的每个标注者ID,其能力值为介于0(随机标注)与1之间的数值,格式为制表符分隔的标注者ID与能力值。 音频文件可从https://zenodo.org/record/2589280下载,其使用需遵循对应授权协议。



