遇见数据集

InsectSet47 & InsectSet66: Expanded datasets for automatic acoustic identification of insects (Orthoptera and Cicadidae)

收藏
Zenodo2024-11-18 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

<strong>Updated full version with training, validation and test sets.</strong> Two newly compiled datasets for training neural networks to automatically identify insect species while comparing adaptive, waveform-based frontends to conventional mel-spectrogram frontends for audio feature extraction. This work was published in PLOS Computational Biology and the machine learning implementations were published on Github. These datasets expand on the previously published InsectSet32 by including recently published collections of insect recordings by citizen scientists from around the world. Recordings from BioAcoustica, xeno-canto and iNaturalist, as well as private collections by Baudewijn Odé were downloaded and manually inspected. Files with strong noise interference or intense filtering, as well as files containing sounds of multiple species were removed to compile these datasets. The files were standardised to 44.1 kHz mono WAV files ranging in length from less than one second to several minutes. Files containing long periods without insect sounds were edited into multiple smaller files with silent periods no longer than 5 seconds. These files are marked as edits in the annotation file and should be assigned together into train/validation/test sets to prevent data leakage. The annotation files contain information for each recording, including the file name, species name and identifier, as well as the data subset they were included in for training the neural network (training, test, validation). InsectSet47 expands on InsectSet32 with recordings from xeno-canto and contains 1006 original recordings from 47 species, with at least ten files per species. The total length of InsectSet47 is 22 hours. InsectSet66 further expands on InsectSet47 by adding research-grade audio observations from iNaturalist, with a total of 1554 recordings from 66 species, a total length of over 24 hours and a minimum of ten files per species. The datasets were split into the training, validation and test sets while ensuring a roughly equal distribution of audio files and audio material for every species in all three subsets. This resulted in a 60/20/20 split (train/validation/test) by file number and a 64/19.5/16.5 split by file length.

**包含训练集、验证集与测试集的更新完整版数据集**。本研究构建了两套全新数据集,用于训练可自动识别昆虫物种的神经网络,同时对比基于自适应波形的前端与传统梅尔频谱图(mel-spectrogram)前端在音频特征提取中的性能表现。本研究成果已发表于《PLOS计算生物学》,相关机器学习实现代码已开源至GitHub平台。本数据集基于此前发布的InsectSet32进行扩展,纳入了全球各地公民科学家近期录制的昆虫鸣声数据集。研究人员从BioAcoustica、xeno-canto、iNaturalist平台以及Baudewijn Odé的私人收藏库中下载了相关录音,并进行了人工质检。在数据集构建过程中,已剔除存在严重噪声干扰、过度滤波,以及包含多种生物鸣声的音频文件。所有音频文件均统一为44.1kHz采样率的单声道WAV格式,时长区间为不足1秒至数分钟。对于包含长时间无昆虫鸣声片段的文件,研究人员将其剪辑为多个子文件,且各子文件间的静音时长不超过5秒。此类经过剪辑的子文件会在标注文件中标记为“编辑版”,为避免数据泄露,需将同一份原始录音剪辑出的所有子文件统一划分至训练集、验证集或测试集的同一子集内。标注文件包含每条录音的元数据,包括文件名、物种名称与物种标识符,以及该录音所属的神经网络训练子集(训练集、测试集、验证集)。InsectSet47基于InsectSet32扩展,纳入了xeno-canto平台的录音数据,共包含47个物种的1006条原始录音,每个物种至少包含10条录音文件,总时长为22小时。InsectSet66则进一步基于InsectSet47扩展,新增了iNaturalist平台的专业级音频观测数据,共包含66个物种的1554条录音,总时长超过24小时,每个物种至少包含10条录音文件。数据集划分时,确保每个物种的音频文件数量与总音频时长在训练集、验证集、测试集三个子集内的分布大致均衡,最终按文件数量划分的比例为训练集:验证集:测试集=60:20:20,按总音频时长划分的比例为64:19.5:16.5。

提供机构:
Zenodo
创建时间:
2023-08-16
二维码
社区交流群
二维码
科研交流群
商业服务