Frozen Perch 2.0 Audio Embeddings of XenoCanto and iNaturalist with Spatiotemporal Context
收藏资源简介:
This dataset contains precomputed audio embeddings of birdsong audio from nearly 10,000 bird species, with audio samples from XenoCanto and iNaturalist. While both XenoCanto and iNaturalist share their bioacoustic audio and metadata with the public, the raw audio data is very large. Perch 2.0 is a publicly available bioacoustic model designed for extracting frozen embeddings from audio samples, which are designed for training bioacoustic classifiers. Thus, we release this dataset of precomputed, frozen embeddings from two popular bioacoustic training datasets to support lightweight training of bioacoustic classifiers with embeddings from Perch 2.0. These precomputed embeddings include both embeddings of randomly sampled 5-second windows of source recordings, as well as samples that use mixup (a data augmentation strategy where audio from multiple recordings are mixed together to create synthetic multi-label samples that include multiple species vocalizing simultaneously, approximating soundscape data), as well as a location-based mixup variant that mixes audio samples with audio samples from within the 100 nearest geographic neighbors in the dataset (according to recording location), aiming to create geographically plausible soundscapes containing species that live in the same environment. The dataset contains geographic coordinates and temporal metadata for each sample, to support spatiotemporally-aware classification, where classifiers predict species conditioned on both audio and the context of where and when the audio was recorded. Code for how these embeddings are precomputed and for using the embeddings, labels and metadata to train bioacoustic classifiers is available on GitHub. This data also includes pretrained weights for a spatiotemporally-aware, global-scale bird species classifier trained on these embeddings, which can be used with the code from the same GitHub repository.



