遇见数据集

Rana sierrae and muscosa vocalizations 2025

收藏
Zenodo2026-07-27 更新2026-08-01 收录
官方服务:

资源简介:

# Rana sierrae and Rana muscosa vocalizations dataset This dataset contains passive acoustic recordings of underwater soundscapes recorded in the habitats of _Rana sierrae_ and _Rana muscosa_ (the mountain yellow-legged frog species complex) in 2022, 2023, 2024, and 2025. The dataset is associated with a manuscript in preparation, in which we perform comparative analyses to investigate geographic variation in the calls of these endangered species across populations, genetic clades, and species. Site names have been anonymized and coordinates have been jittered to protect these endangered species. Please contact the authors if you need access to the original site names and coordinates. The dataset contains 3 subsets: 1. A large set of 48,763 1-minute soundscape recordings in which _R. sierrae_ / _muscosa_ vocalizations were detected by an automated classifier, and associated metadata (`all_call_detections/`, 21 Gb). We also include a set of acoustic features extracted from (1) 261,270 automatically detected calls, and (2) the subset of 2,511 manually annotated calls. 2. A set of 2,511 4-second audio clips in which _R. sierrae_ / _muscosa_ vocalizations were manually annotated by an expert observer (`verified_call_detections/`, 79 Mb) 3. An evaluation set of 1000 audio clips and associated presence/absence labels used to evaluate machine learning classifier performance (`annotated_test_set/`, 18 Mb) ## Usage Each of the subfolders is provided as a .zip file for download. Unzip the files into the top-level dataset directory to access and analyze the audio files. Because `all_call_detections/1m_files/` is large (~20 Gb, 48,763 files), it is instead split across 20 approximately even-sized shard zips (`1m_files_shard_00.zip` ... `1m_files_shard_19.zip`), plus a separate `all_call_detections_csvs.zip` containing the csv files from `all_call_detections/`. Both are provided at the top level of the dataset directory alongside the other zip files. To unzip these into place, download all 21 zip files (the 20 shards + `all_call_detections_csvs.zip`) into the top-level dataset directory and run `unshard_unzip_1m_files.py` from that directory: ``` python3 unshard_unzip_1m_files.py ``` This extracts each shard's mp3 files into `all_call_detections/1m_files/` and the csv files into `all_call_detections/`, recreating the full `all_call_detections/` folder structure shown below. (`shard_zip_1m_files.py` is the corresponding script used to create the shard and csv zips in the first place, and does not need to be run by users of the dataset.) Scripts reproducing the automated detection, feature extraction, and other analyses are available in the GitHub repository: `https://github.com/sammlapp/rana_sierrae_comparative_bioacoustics/` ## Methodology These data were created by: 1. Collecting passive acoustic monitoring data using underwater AudioMoth recorders 2. Developing then applying a binary classifier CNN and an object detecting CNN to locate vocalizations in the full dataset 3. Performing manual review of random clips and detected clips to create the set of manually verified detections and the evaluation set 4. Applying custom acoustic feature extraction scripts See the README.md for further details

提供机构:
Zenodo
创建时间:
2026-07-27
二维码
社区交流群
二维码
科研交流群
商业服务