遇见数据集

SEA-BIRD: a machine learning–ready dataset for Malaysian garden bird sounds

收藏
Zenodo2026-05-20 更新2026-05-26 收录
官方服务:

资源简介:

The SEA-Bird dataset is a curated, machine-learning–ready collection of avian vocalizations from ten bird species that are among the most commonly found in Malaysia. Recordings are sourced primarily from Southeast Asia, with a small number from surrounding regions. The dataset comprises 6,000 three-second audio clips sampled at 16 kHz and an additional 5,792 clips at 44.1 kHz, each manually verified to ensure minimal background noise, balanced class representation, and the absence of overlapping species calls. The SEA-Bird dataset was developed to address the scarcity of machine-learning–ready acoustic datasets from underrepresented tropical regions, particularly in Southeast Asia. It supports research in automated species identification, biodiversity monitoring, and embedded acoustic sensing. The 16 kHz subset is optimized for Edge AI applications on low-power microcontrollers, while the 44.1 kHz version preserves high-frequency detail suitable for advanced spectro-temporal analyses. All recordings were sourced from the Xeno-canto open repository under Creative Commons licenses and subjected to a rigorous curation process. This process included segmentation and manual inspection of spectrograms to ensure vocal clarity and completeness. The dataset is organized into training, validation, and test directories. Each directory contains ten subfolders, one for each bird species included in the dataset (e.g., Common_Tailorbird/, Spotted_Dove/). Within each species folder, individual clips are named using the original Xeno-canto recording ID followed by the clip start time in milliseconds relative to the source recording (e.g., XC123456_15000.wav). This naming and directory structure preserves traceability and facilitates seamless integration into automated processing pipelines. The training, validation, and test splits follow an exact 75/10/15 ratio and were generated using mixed-integer programming to eliminate data leakage. No source recording contributes clips to more than one split. The split ratios were chosen to optimize the size of the training split while keeping enough samples for validation and testing. Preliminary experiments using MobileNetV3-Small achieved classification accuracies above 85%, while benchmark models such as EfficientNet-B0, ResNet-50, and VGG-16 confirmed the dataset’s robustness for deep-learning applications. Both subsets are released under a Creative Commons Attribution (CC BY 4.0) license and are freely available for academic and applied research. The dataset contributes to the growing body of open bioacoustic resources that promote reproducible, scalable, and inclusive research in tropical biodiversity informatics.

提供机构:
Zenodo
创建时间:
2026-02-20
二维码
社区交流群
二维码
科研交流群
商业服务