遇见数据集

sounding_out_chorus

收藏
Zenodo2023-12-20 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains the data for the paper Towards interpretable learned representations for Ecoacoustics using variational auto-encoding. This dataset contains a series of 1 min wav files recorded across UK and Ecuadorian habitats. Each sample has 26 acoustic indices calculated and a full list of avian species and abundances. This dataset is an updated version of a previous release (10.5281/zenodo.1255218) including KML maps containing GPS data for each sample site and updated label metadata. A data module for use in a PyTorch machine learning pipeline is available here. Abstract Ecoacoustics is an emerging science that seeks to understand the role of sound in ecological processes. Passive acoustic monitoring is increasingly being used to collect vast quantities of whole-soundscape audio recordings in order to study variations in acoustic community activity across spatial and temporal scales. However, extracting relevant information from audio recordings for ecological inference is non-trivial. Recent approaches to machine-learned acoustic features appear promising but are limited by inductive biases, crude temporal integration methods and few means to interpret downstream inference. To address these limitations we developed and trained a self-supervised representation learning algorithm - a convolutional Variational Auto-Encoder (VAE) - to embed latent features from acoustic survey data collected from sites representing a gradient of habitat degradation in temperate and tropical ecozones and use prediction of survey site as a test case for interpreting inference. We investigate approaches to interpretability by mapping discriminative descriptors back to the spectro-temporal domain to observe how soundscape components change as we interpolate across a linear classification boundary traversing latent feature space; we advance temporal integration methods by encoding a probabilistic soundscape descriptor capable of capturing multi-modal distributions of latent features over time. Our results suggest that varying combinations of soundscape components (biophony, geophony and anthrophony) are used to infer sites along a degradation gradient and increased sensitivity to periodic signals improves on previous research using time-averaged representations for site classification. We also find the VAE is highly sensitive to differences in recorder hardware’s frequency response and demonstrate a simple linear transformation to mitigate the effect of hardware variance on the learned representation. Our work paves the way for development of a new class of deep neural networks that afford more interpretable machine-learned ecoacoustic representations to advance the fundamental and applied science and support global conservation efforts. Sampling Methods (extract from paper) Surveys were designed to monitor the acoustic characteristics of sites across a gradient of degradation, ranging from primary forest, through secondary forest (or areas in the process of ecological restoration), to agricultural monocultures, providing a space-for-time substitution to investigate changes in soundscapes across a gradient of ecological status. Samples were taken for 1 minute in every 15 for 10 sequential days at each site. Full dawn and dusk recordings were also collected. In each site, 15 recorders were placed in a grid-like system spaced a minimum of 200m away from their neighbours in the UK - 300m in Ecuador - to mitigate acoustic overlap and avoid spatial pseudo-replication. Wildlife Acoustics Song Meters equipped with two channel omni-directional microphone were used. Seven SM2+ and eight SM3 devices were deployed. Gains were matched between recorders (analogue gains at +36dB on SM2+ and +12dB on SM3 which has inbuilt +12dB gain) and recordings made at resolution of 16 bits with a sampling rate of 48 kHz. To provide a cleaner validation data set, local weather recordings were used to select 3 days with lowest wind and rain from each site giving 4725 1 min recording in total. Sites are labelled by their quality in descending order i.e. UK1 (primary), UK2 (regenerating), UK3 (degraded).

本仓库收录了论文《面向生态声学的可解释学习表征:基于变分自编码方法》(Towards interpretable learned representations for Ecoacoustics using variational auto-encoding)的配套数据集。本数据集包含一系列时长1分钟的WAV音频文件,采集自英国与厄瓜多尔的各类生境。每份样本均计算了26项声学指数,并附有鸟类物种与种群丰度的完整列表。本数据集为此前发布版本(DOI: 10.5281/zenodo.1255218)的更新版,新增了包含各采样点位GPS数据的KML地图,以及更新后的标签元数据。适用于PyTorch机器学习流水线的数据模块已在此处提供。 摘要 生态声学是一门新兴学科,旨在探究声音在生态过程中的作用。被动声学监测正愈发广泛地用于采集海量全声景音频录音,以研究声学群落活动在空间与时间尺度上的变化。然而,从音频录音中提取可用于生态推断的有效信息并非易事。近年来,基于机器学习的声学特征方法展现出应用前景,但仍受限于归纳偏置、粗糙的时间整合手段,以及下游推断可解释性手段匮乏等问题。为解决上述局限,我们开发并训练了一种自监督表征学习算法——卷积变分自编码器(convolutional Variational Auto-Encoder, VAE),从温带与热带生态区不同生境退化梯度的采样点采集的声学调查数据中嵌入潜在特征,并以采样点分类作为测试案例,验证推断的可解释性。我们通过将判别性描述符映射到时频域,以观察在穿越潜在特征空间的线性分类边界插值时声景组分的变化,以此探究可解释性方法;同时我们改进了时间整合手段,通过编码一种概率声景描述符,实现对随时间变化的潜在特征多模态分布的捕捉。研究结果表明,声景组分(生物声、地质声与人为声)的不同组合可用于推断不同退化梯度下的采样点,且对周期信号的增强敏感性优于此前基于时间平均表征的点位分类研究。我们还发现,变分自编码器对录音设备的频率响应差异具有高度敏感性,并提出了一种简单的线性变换方法,以缓解硬件差异对学习表征的影响。本研究为新一代深度神经网络的开发铺平了道路,这类网络可提供更具可解释性的机器学习生态声学表征,以推动基础与应用科学研究,并助力全球保护工作。 采样方法(摘自论文) 本调查旨在监测不同退化梯度点位的声学特征,采样梯度涵盖原始森林、次生林(或生态恢复中区域)至农业单作区,通过空间替代时间的方式,探究不同生态状态下声景的变化。每个点位在连续10天内,每15分钟采集1分钟的音频样本,同时还采集了完整的黎明与黄昏时段录音。每个点位部署15台记录仪,在英国采用网格状布局,相邻记录仪间距不低于200米,厄瓜多尔则为300米,以避免声学重叠与空间伪重复。本研究使用配备双通道全向麦克风的Wildlife Acoustics Song Meters记录仪,其中包括7台SM2+与8台SM3设备。记录仪增益设置保持一致:SM2+的模拟增益为+36dB,SM3自带+12dB增益,因此其模拟增益设为+12dB。录音分辨率为16比特,采样率为48kHz。为获得更纯净的验证数据集,我们利用当地气象记录,从每个点位的10天录音中筛选出风速与降雨量最低的3天,最终总计得到4725份1分钟时长的录音。点位按质量从高到低标注,即UK1(原始林)、UK2(恢复中)、UK3(退化生境)。

提供机构:
Zenodo
创建时间:
2023-09-07
二维码
社区交流群
二维码
科研交流群
商业服务