Data from: Distributed acoustic cues for caller identity in macaque vocalization
收藏资源简介:
Individual primates can be identified by the sound of their voice. Macaques have demonstrated an ability to discern conspecific identity from a harmonically structured ‘coo’ call. Voice recognition presumably requires the integrated perception of multiple acoustic features. However, it is unclear how this is achieved, given considerable variability across utterances. Specifically, the extent to which information about caller identity is distributed across multiple features remains elusive. We examined these issues by recording and analysing a large sample of calls from eight macaques. Single acoustic features, including fundamental frequency, duration and Weiner entropy, were informative but unreliable for the statistical classification of caller identity. A combination of multiple features, however, allowed for highly accurate caller identification. A regularized classifier that learned to identify callers from the modulation power spectrum of calls found that specific regions of spectral–temporal modulation were informative for caller identification. These ranges are related to acoustic features such as the call’s fundamental frequency and FM sweep direction. We further found that the low-frequency spectrotemporal modulation component contained an indexical cue of the caller body size. Thus, cues for caller identity are distributed across identifiable spectrotemporal components corresponding to laryngeal and supralaryngeal components of vocalizations, and the integration of those cues can enable highly reliable caller identification. Our results demonstrate a clear acoustic basis by which individual macaque vocalizations can be recognized.
个体灵长类动物可通过其发声实现识别。猕猴(macaques)已被证实能够从具有谐波结构的“咕咕”叫声中辨别同类个体的身份。发声识别理论上需要对多种声学特征进行整合感知。然而,考虑到不同发声间存在显著差异,该识别机制的具体实现方式尚不明确。具体而言,发声者身份信息在多种声学特征间的分布程度仍有待厘清。本研究通过记录并分析8只猕猴的大量叫声样本,对上述问题展开了研究。单声学特征(包括基频(fundamental frequency)、时长以及维纳熵(Weiner entropy))虽可为发声者身份的统计分类提供有效信息,但可靠性不足。然而,将多种特征相结合,则可实现高精度的发声者身份识别。基于叫声调制功率谱(modulation power spectrum)训练得到的正则化分类器(regularized classifier)显示,特定的频谱-时间调制(spectral–temporal modulation)区域可为发声者身份识别提供有效信息。这些调制范围与叫声的基频、调频扫频(FM sweep)方向等声学特征相关。本研究还发现,低频频谱-时间调制成分包含了发声者体型的指示性线索。综上,发声者身份的线索分布于可识别的频谱-时间调制成分中,这些成分对应发声过程中的喉部(laryngeal)与喉上(supralaryngeal)结构,而对这些线索的整合可实现高度可靠的发声者身份识别。本研究结果明确了猕猴个体叫声可被识别的声学基础。




