遇见数据集

SEA-BIRD: a machine learning–ready dataset for Malaysian garden bird sounds

收藏
Zenodo2026-05-20 更新2026-05-26 收录
官方服务:

资源简介:

The SEA-Bird dataset is a curated, machine-learning–ready collection of avian vocalizations from ten bird species that are among the most commonly found in Malaysia. Recordings are sourced primarily from Southeast Asia, with a small number from surrounding regions. The dataset comprises 6,000 three-second audio clips sampled at 16 kHz and an additional 5,792 clips at 44.1 kHz, each manually verified to ensure minimal background noise, balanced class representation, and the absence of overlapping species calls. The SEA-Bird dataset was developed to address the scarcity of machine-learning–ready acoustic datasets from underrepresented tropical regions, particularly in Southeast Asia. It supports research in automated species identification, biodiversity monitoring, and embedded acoustic sensing. The 16 kHz subset is optimized for Edge AI applications on low-power microcontrollers, while the 44.1 kHz version preserves high-frequency detail suitable for advanced spectro-temporal analyses. All recordings were sourced from the Xeno-canto open repository under Creative Commons licenses and subjected to a rigorous curation process. This process included segmentation and manual inspection of spectrograms to ensure vocal clarity and completeness. The dataset is organized into training, validation, and test directories. Each directory contains ten subfolders, one for each bird species included in the dataset (e.g., Common_Tailorbird/, Spotted_Dove/). Within each species folder, individual clips are named using the original Xeno-canto recording ID followed by the clip start time in milliseconds relative to the source recording (e.g., XC123456_15000.wav). This naming and directory structure preserves traceability and facilitates seamless integration into automated processing pipelines. The training, validation, and test splits follow an exact 75/10/15 ratio and were generated using mixed-integer programming to eliminate data leakage. No source recording contributes clips to more than one split. The split ratios were chosen to optimize the size of the training split while keeping enough samples for validation and testing. Preliminary experiments using MobileNetV3-Small achieved classification accuracies above 85%, while benchmark models such as EfficientNet-B0, ResNet-50, and VGG-16 confirmed the dataset’s robustness for deep-learning applications. Both subsets are released under a Creative Commons Attribution (CC BY 4.0) license and are freely available for academic and applied research. The dataset contributes to the growing body of open bioacoustic resources that promote reproducible, scalable, and inclusive research in tropical biodiversity informatics.

SEA-Bird数据集(SEA-Bird dataset)是一套经过精心甄选整理、可直接用于机器学习的鸟类鸣声数据集,收录了马来西亚境内最为常见的10种鸟类的鸣声。该数据集的录音素材主要采集自东南亚地区,少量来自周边区域。数据集包含6000段时长为3秒、采样率为16kHz的音频片段,以及额外5792段采样率为44.1kHz的同类片段;所有片段均经过人工校验,以确保背景噪音极低、类别分布均衡且无跨物种鸣声重叠的情况。 SEA-Bird数据集的研发初衷,是为了解决热带地区(尤其是东南亚)缺乏可直接用于机器学习的声学数据集这一痛点。该数据集可支撑鸟类物种自动识别、生物多样性监测以及嵌入式声学传感等方向的研究。其中16kHz子集针对低功耗微控制器上的边缘人工智能(Edge AI)应用做了优化,而44.1kHz版本则保留了高频细节,适用于高级频谱时序分析。 所有录音素材均来自Xeno-canto开源库,遵循知识共享(Creative Commons)许可协议,并经过了严格的整理流程。该流程包括对频谱图(spectrogram)进行分段与人工检视,以确保鸣声清晰且完整。 数据集按照训练集、验证集与测试集划分至不同目录中。每个目录下设10个子文件夹,分别对应数据集中的10种鸟类(例如:Common_Tailorbird/、Spotted_Dove/)。在每个鸟类物种的子文件夹中,单条音频片段的命名格式为:原Xeno-canto录音ID加上该片段相对于源录音的起始时间(单位为毫秒),例如XC123456_15000.wav。该命名与目录结构可保证数据溯源性,并能无缝集成至自动化处理流水线中。 训练集、验证集与测试集的划分比例严格遵循75:10:15,且通过混合整数规划生成,以避免数据泄露(data leakage)。任何一条源录音的片段均不会被分配至多个划分集中。该划分比例的设定旨在最大化训练集规模,同时保留足够的样本用于验证与测试。 使用MobileNetV3-Small进行的初步实验已取得85%以上的分类准确率,而EfficientNet-B0、ResNet-50与VGG-16等基准模型的测试结果也证实,该数据集在深度学习应用中具备良好的鲁棒性。 两个子集均采用知识共享署名4.0(Creative Commons Attribution (CC BY 4.0))许可协议发布,可免费用于学术与应用研究。该数据集为不断壮大的开源生物声学资源库添砖加瓦,有助于推动热带生物多样性信息学领域可复现、可扩展且兼具包容性的研究工作。

提供机构:
Zenodo
创建时间:
2025-11-25
二维码
社区交流群
二维码
科研交流群
商业服务