遇见数据集

Acoustic recordings of three Hainan gibbon (Nomascus hainanus) social groups for machine learning and deep learning analyses

收藏
Zenodo2026-06-24 更新2026-06-21 收录
官方服务:

资源简介:

This dataset is released as a companion resource to the paper “Passive acoustic identification of social groups in the Hainan gibbon” (DOI: https://doi.org/10.1002/rse2.70088). It contains the acoustic recordings used in the study and is provided to enable reproducible benchmarking and methodological development in bioacoustic machine learning. The recordings were collected from three distinct Hainan gibbon (Nomascus hainanus) social groups in natural habitat conditions and are intended for use in machine learning and deep learning applications, including social group identification, bioacoustic classification, and ecological monitoring. The dataset includes raw audio files, corresponding phrase-level annotations, and segment-level annotations. In addition, it provides predefined training, validation, and test partitions used in the original study to support reproducible training and evaluation of segment- and sequence-based models. Files provided: Raw audio recordings The five compressed archives, part_1–5.zip, contain 97 eight-hour raw audio recordings collected from three Hainan gibbon social groups: Group B, Group C, and Group D. The recordings are distributed as follows: Group B: 36 recordings Group C: 24 recordings Group D: 37 recordings Annotations The compressed archive Phrase_and_segment_annotations.zip contains the following CSV files: phrase_annotations.csv This file contains the start and end times (in seconds) of each annotated phrase within the corresponding audio recording. In total, the dataset includes 7,958 annotated phrases, distributed as follows: Group B: 2,630 phrases Group C: 2,932 phrases Group D: 2,396 phrases all_segment_annotations.csv This file contains 32,706 non-overlapping 1-second segments generated from the annotated phrases. The segments are distributed as follows: Group B: 10,259 segments Group C: 11,575 segments Group D: 10,872 segments train_segment_annotations.csv, validation_segment_annotations.csv, and test_segment_annotations.csv These files contain the segments used to construct the training, validation, and test datasets from which segment-based and sequence-based samples are generated. Each file includes the following information: Segment identifier Segment start and end sample indices within the source audio recording Source audio recording filename Social group label associated with the recording Closed-Population Identification Data The compressed archive closed_population_ID.zip contains the training, validation, and test datasets used for segment-based and sequence-based classification experiments under a closed-population identification setting, where the set of Hainan gibbon social groups is assumed to be known and fixed throughout the study period. Open-Population Identification Data The compressed archive open_population_ID.zip contains pair-based segment and sequence datasets used for similarity learning experiments under an open-population identification setting, where the number of Hainan gibbon social groups is assumed to be unknown and may change over time. Implementation Details Detailed information on phrase extraction, segment generation, sequence construction, spectrogram and MFCC computation, model development, training procedures, and evaluation protocols is available in the accompanying GitHub-Zenodo software archive: Software archive: DOI: 10.5281/zenodo.20832289 Users wishing to reproduce the experiments reported in "Passive acoustic identification of social groups in the Hainan gibbon" should consult the software archive in conjunction with this dataset. Citation Users of this dataset are encourage to cite both the dataset and its associate paper as listed above.

提供机构:
Zenodo
创建时间:
2026-06-18
二维码
社区交流群
二维码
科研交流群
商业服务