遇见数据集

Automatic Recognition of Element Classes and Boundaries in the Birdsong with Variable Sequences

收藏
Figshare2016-09-28 更新2026-04-29 收录
官方服务:

资源简介:

Researches on sequential vocalization often require analysis of vocalizations in long continuous sounds. In such studies as developmental ones or studies across generations in which days or months of vocalizations must be analyzed, methods for automatic recognition would be strongly desired. Although methods for automatic speech recognition for application purposes have been intensively studied, blindly applying them for biological purposes may not be an optimal solution. This is because, unlike human speech recognition, analysis of sequential vocalizations often requires accurate extraction of timing information. In the present study we propose automated systems suitable for recognizing birdsong, one of the most intensively investigated sequential vocalizations, focusing on the three properties of the birdsong. First, a song is a sequence of vocal elements, called notes, which can be grouped into categories. Second, temporal structure of birdsong is precisely controlled, meaning that temporal information is important in song analysis. Finally, notes are produced according to certain probabilistic rules, which may facilitate the accurate song recognition. We divided the procedure of song recognition into three sub-steps: local classification, boundary detection, and global sequencing, each of which corresponds to each of the three properties of birdsong. We compared the performances of several different ways to arrange these three steps. As results, we demonstrated a hybrid model of a deep convolutional neural network and a hidden Markov model was effective. We propose suitable arrangements of methods according to whether accurate boundary detection is needed. Also we designed the new measure to jointly evaluate the accuracy of note classification and boundary detection. Our methods should be applicable, with small modification and tuning, to the songs in other species that hold the three properties of the sequential vocalization.

针对序列发声的研究往往需要对长时连续声学信号中的发声行为进行分析。在诸如发育研究或跨世代研究这类需要分析数日乃至数月发声行为的场景中,自动识别方法是学界迫切需要的技术手段。尽管面向应用场景的自动语音识别(Automatic Speech Recognition, ASR)方法已得到大量深入研究,但直接将其套用于生物学研究未必是最优方案。这是因为与人类语音识别不同,序列发声分析往往需要精准提取时序信息。本研究以鸣禽鸣唱——目前研究最为深入的序列发声类型之一——为对象,针对其三大特性提出了适配的自动化识别系统:其一,鸣唱由可被归类为不同类别的发声单元(note)序列构成;其二,鸣禽鸣唱的时序结构受到精准调控,这意味着时序信息在鸣唱分析中至关重要;其三,音节的产生遵循特定的概率规则,这有助于实现精准的鸣唱识别。我们将鸣唱识别流程拆分为三个子步骤:局部分类、边界检测与全局序列建模,三者分别对应鸣禽鸣唱的上述三大特性。我们对比了多种不同的子步骤排列组合方案的识别性能,实验结果表明,深度卷积神经网络(Deep Convolutional Neural Network, DCNN)与隐马尔可夫模型(Hidden Markov Model, HMM)的混合模型具备优异的识别效果。我们针对是否需要精准边界检测的场景,提出了适配的方法组合方案;此外,我们设计了一种全新的评估指标,可同时衡量音节分类与边界检测的精度。本研究提出的方法仅需少量修改与调参,即可推广应用于具备序列发声三大特性的其他物种的鸣唱分析任务。

创建时间:
2016-09-28
二维码
社区交流群
二维码
科研交流群
商业服务