遇见数据集

SEA-Bird: A Machine Learning–Ready Bird Sound Dataset for Ten Common Southeast Asian Species

收藏
Zenodo2026-05-08 更新2026-05-26 收录
官方服务:

资源简介:

The SEA-Bird dataset is a curated, machine-learning–ready collection of avian vocalizations from ten bird species that are among the most commonly found in Malaysia. Recordings are sourced primarily from Southeast Asia, with a small number from surrounding regions. The dataset comprises 6,000 three-second audio clips sampled at 16 kHz and an additional 5,821 clips at 44.1 kHz, each manually verified to ensure minimal background noise, balanced class representation, and the absence of overlapping species calls. The SEA-Bird dataset was developed to address the scarcity of machine-learning–ready acoustic datasets from underrepresented tropical regions, particularly in Southeast Asia. It supports research in automated species identification, biodiversity monitoring, and embedded acoustic sensing. The 16 kHz subset is optimized for Edge AI applications on low-power microcontrollers, while the 44.1 kHz version preserves high-frequency detail suitable for advanced spectro-temporal analyses. All recordings were obtained from the Xeno-canto open repository under Creative Commons licenses and underwent a rigorous curation process involving segmentation and manual spectrogram inspection to ensure vocal clarity and completeness. The dataset is organized into ten folders, each named after one of the included bird species (e.g., Common_Myna/, Spotted_Dove/). Within each folder, individual clips are labeled using the original Xeno-canto recording ID, followed by the start time in milliseconds from the source recording (e.g., XC123456_15000.wav). This structure preserves traceability and supports easy integration into automated pipelines. This release does not yet include separate metadata files; a future update will provide a CSV metadata table linking each clip to its source recording and associated attributes. Preliminary experiments using MobileNetV3-Small achieved classification accuracies above 90%, while benchmark models such as EfficientNet-B0, ResNet-50, and VGG-16 confirmed the dataset’s robustness for deep-learning applications. Both subsets are released under a Creative Commons Attribution (CC BY 4.0) license and are freely available for academic and applied research. The dataset contributes to the growing body of open bioacoustic resources that promote reproducible, scalable, and inclusive research in tropical biodiversity informatics.

提供机构:
Zenodo
创建时间:
2025-11-25
二维码
社区交流群
二维码
科研交流群
商业服务