遇见数据集

ARCA23K

收藏
Zenodo2022-02-25 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

ARCA23K is a dataset of labelled sound events created to investigate real-world label noise. It contains 23,727 audio clips originating from Freesound, and each clip belongs to one of 70 classes taken from the AudioSet ontology. The dataset was created using an entirely automated process with no manual verification of the data. For this reason, many clips are expected to be labelled incorrectly. In addition to ARCA23K, this release includes a companion dataset called ARCA23K-FSD, which is a single-label subset of the FSD50K dataset. ARCA23K-FSD contains the same sound classes as ARCA23K and the same number of audio clips per class. As it is a subset of FSD50K, each clip and its label have been manually verified. Note that only the ground truth data of ARCA23K-FSD is distributed in this release. To download the audio clips, please visit the Zenodo page for FSD50K. A paper has been published detailing how the dataset was constructed. See the Citing section below. The source code used to create the datasets is available: https://github.com/tqbl/arca23k-dataset <strong>Characteristics</strong> ARCA23K(-FSD) is divided into: A training set containing 17,979 clips (39.6 hours for ARCA23K). A validation set containing 2,264 clips (5.0 hours). A test test containing 3,484 clips (7.3 hours). There are 70 sound classes in total. Each class belongs to the AudioSet ontology. Each audio clip was sourced from the Freesound database. Other than format conversions (e.g. resampling), the audio clips have not been modified. The duration of the audio clips varies from 0.3 seconds to 30 seconds. All audio clips are mono 16-bit WAV files sampled at 44.1 kHz. Based on listening tests (details in paper), 46.4% of the training examples are estimated to be labelled incorrectly. Among the incorrectly-labelled examples, 75.9% are estimated to be out-of-vocabulary. <strong>Sound Classes</strong> The list of sound classes is given below. They are grouped based on the top-level superclasses of the AudioSet ontology. <em>Music</em> Acoustic guitar Bass guitar Bowed string instrument Crash cymbal Electric guitar Gong Harp Organ Piano Rattle (instrument) Scratching (performance technique) Snare drum Trumpet Wind chime Wind instrument, woodwind instrument <em>Sounds of things</em> Boom Camera Coin (dropping) Computer keyboard Crack Dishes, pots, and pans Drawer open or close Drill Gunshot, gunfire Hammer Keys jangling Knock Microwave oven Printer Sawing Scissors Skateboard Slam Splash, splatter Squeak Tap Thump, thud Toilet flush Train Water tap, faucet Whoosh, swoosh, swish Writing Zipper (clothing) <em>Natural sounds</em> Crackle Stream Waves, surf Wind <em>Human sounds</em> Burping, eructation Chewing, mastication Child speech, kid speaking Clapping Cough Crying, sobbing Fart Female singing Female speech, woman speaking Finger snapping Giggle Male speech, man speaking Run Screaming Walk, footsteps <em>Animal</em> Bark Cricket Livestock, farm animals, working animals Meow Rattle <em>Source-ambiguous sounds</em> Crumpling, crinkling Crushing Tearing <strong>License and Attribution</strong> This release is licensed under the Creative Commons Attribution 4.0 International License. The audio clips distributed as part of ARCA23K were sourced from Freesound and have their own Creative Commons license. The license information and attribution for each audio clip can be found in <code>ARCA23K.metadata/train.json</code>, which also includes the original Freesound URLs. The files under <code>ARCA23K-FSD.ground_truth/</code> are an adaptation of the ground truth data provided as part of FSD50K, which is licensed under the Creative Commons Attribution 4.0 International License. The curators of FSD50K are Eduardo Fonseca, Xavier Favory, Jordi Pons, Mercedes Collado, Ceren Can, Rachit Gupta, Javier Arredondo, Gary Avendano, and Sara Fernandez. <strong>Citing</strong> If you wish to cite this work, please cite the following paper: T. Iqbal, Y. Cao, A. Bailey, M. D. Plumbley, and W. Wang, “ARCA23K: An audio dataset for investigating open-set label noise”, in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021), 2021, Barcelona, Spain, pp. 201–205. BibTeX: <pre>@inproceedings{Iqbal2021, author = {Iqbal, T. and Cao, Y. and Bailey, A. and Plumbley, M. D. and Wang, W.}, title = {{ARCA23K}: An audio dataset for investigating open-set label noise}, booktitle = {Proceedings of the Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021)}, pages = {201--205}, year = {2021}, address = {Barcelona, Spain}, }</pre>

ARCA23K是专为探究现实世界标签噪声而构建的标注音频事件数据集。该数据集包含23727段源自Freesound平台(Freesound)的音频片段,每段音频隶属于从AudioSet本体(AudioSet ontology)中选取的70个类别之一。本数据集完全通过自动化流程构建,未经过人工数据验证,因此预计存在大量标注错误的片段。 除ARCA23K外,本次发布还包含配套数据集ARCA23K-FSD,它是FSD50K数据集(FSD50K)的单标注子集。ARCA23K-FSD与ARCA23K拥有相同的声音类别,且每个类别的音频片段数量一致。由于其作为FSD50K的子集,每段音频及其标注均经过人工验证。请注意,本次发布仅分发ARCA23K-FSD的真实标注数据,若需下载音频片段,请访问FSD50K的Zenodo页面(Zenodo)。已有论文详细阐述了该数据集的构建方式,详见下文的引用部分。用于构建数据集的源代码可在以下链接获取:https://github.com/tqbl/arca23k-dataset ### 数据集特性 ARCA23K(含ARCA23K-FSD)被划分为: - 训练集:包含17979段音频片段(ARCA23K版本总时长39.6小时) - 验证集:包含2264段音频片段(总时长5.0小时) - 测试集:包含3484段音频片段(总时长7.3小时) 总计包含70个声音类别,均隶属于AudioSet本体。所有音频片段均源自Freesound平台。除格式转换(如重采样)外,音频片段未经过其他修改。音频片段的时长范围为0.3秒至30秒,全部为单声道16位WAV文件,采样率为44.1kHz。根据听音测试(详细内容见论文),估计有46.4%的训练样本存在标注错误,其中75.9%的错误标注样本属于未登录词(out-of-vocabulary)范畴。 ### 声音类别 下文列出了所有声音类别,它们按照AudioSet本体的顶层超类进行分组: #### 音乐类 原声吉他、贝斯吉他、弓弦乐器、碎音钹(crash cymbal)、电吉他、锣、竖琴、管风琴、钢琴、摇奏乐器(rattle (instrument))、刮奏(演奏技法,scratching (performance technique))、小军鼓、小号、风铃、木管乐器 #### 物体声响类 爆炸声、相机声响、硬币掉落声、电脑键盘声、碎裂声、锅碗瓢盆声、抽屉开合声、钻孔声、枪声、锤击声、钥匙碰撞声、敲击声、微波炉声响、打印机声响、锯切声、剪刀声、滑板滑行声、撞击声、飞溅声、吱吱声、轻拍声、重击声、马桶冲水声、火车声响、水龙头流水声、呼啸声、书写声、拉链声 #### 自然声响类 噼啪声、溪流声、海浪拍击声、风声 #### 人类声响类 打嗝声、咀嚼声、儿童言语声、拍手声、咳嗽声、哭泣声、放屁声、女性歌唱声、女性言语声、打响指声、咯咯笑声、男性言语声、跑步声、尖叫声、行走脚步声 #### 动物声响类 犬吠声、蟋蟀鸣叫声、家畜(农场动物、役用动物)声响、猫叫声、响尾声(rattle) #### 源歧义声响类 揉皱声、压碎声、撕裂声 ### 许可与归因 本次发布采用知识共享署名4.0国际许可协议(Creative Commons Attribution 4.0 International License)进行授权。作为ARCA23K组成部分分发的音频片段源自Freesound平台,各自带有对应的知识共享许可。每段音频片段的许可信息与归因详情可在`ARCA23K.metadata/train.json`文件中查看,该文件同时包含原始Freesound链接。`ARCA23K-FSD.ground_truth/`目录下的文件是对FSD50K提供的真实标注数据的改编,FSD50K采用知识共享署名4.0国际许可协议进行授权。FSD50K的策展人为Eduardo Fonseca、Xavier Favory、Jordi Pons、Mercedes Collado、Ceren Can、Rachit Gupta、Javier Arredondo、Gary Avendano与Sara Fernandez。 ### 引用方式 若需引用本研究,请引用以下论文:T. Iqbal、Y. Cao、A. Bailey、M. D. Plumbley与W. Wang,“ARCA23K:用于探究开放集标签噪声的音频数据集”,收录于《2021年声学场景与事件检测与分类研讨会论文集(DCASE2021)》,2021年,西班牙巴塞罗那,第201-205页。BibTeX格式如下: @inproceedings{Iqbal2021, author = {Iqbal, T. and Cao, Y. and Bailey, A. and Plumbley, M. D. and Wang, W.}, title = {{ARCA23K}: An audio dataset for investigating open-set label noise}, booktitle = {Proceedings of the Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021)}, pages = {201--205}, year = {2021}, address = {Barcelona, Spain}, }

提供机构:
Zenodo
创建时间:
2021-09-15
二维码
社区交流群
二维码
科研交流群
商业服务