Finding, visualizing, and quantifying latent structure across diverse animal vocal repertoires.

Finding, visualizing, and quantifying latent structure across diverse animal vocal repertoires.
复制标题

发现,可视化和量化各种动物声曲目的潜在结构。

DOI:
10.1371/journal.pcbi.1008228
复制
发表时间:
2020-10
影响因子:
4.3
通讯作者:
Gentner TQ
Gentner TQ
中科院分区:
生物学2区
文献类型:
--
作者:
Sainburg T;Thielk M;Gentner TQ

文献摘要

参考文献

相似文献

动物发出的声音很复杂,从一个重复的叫声到数百个独特的声音元素,在几个小时内依次展开。描述复杂的发声需要相当大的努力和对每个物种发声行为的深刻直觉。即使拥有丰富的经验,人类对动物交流的描述也会受到人类感知偏见的影响。我们提出了一套计算方法,用于将动物发声投射到低维潜在表征空间中,这些空间直接从声音信号的频谱图中学习。我们将这些方法应用于来自20多个物种的不同数据集,包括人类,蝙蝠,鸣禽,小鼠,鲸目动物和非人灵长类动物。潜在投影以直观和可量化的方式揭示数据的复杂特征,使声乐声学的高功率比较分析成为可能。我们介绍了作为离散序列和连续潜在变量分析发声的方法。每种方法都可以用于解开复杂的光谱-时间结构和观察通信中的长时间组织。在数千种用声音交流的物种中,只有极少数的声音被描述或详细研究过。这在很大程度上是由于传统的分析方法需要高水平的专业知识,很难开发,而且往往是物种特异性的。在这里,我们提出了一套无监督的方法,将动物发声投射到潜在特征空间中,以定量比较和发展动物发声的视觉直觉。我们通过对来自29个不同物种(包括鸣禽、小鼠、猴子、人类和鲸鱼)的19个动物发声数据集的一系列分析来证明这些方法。我们展示了学习的潜在特征空间如何解开复杂的光谱-时间结构,实现跨物种比较,并揭示发声的高级属性,如声音元素集群的刻板印象,人口区域,协同发音和个体身份。
Animals produce vocalizations that range in complexity from a single repeated call to hundreds of unique vocal elements patterned in sequences unfolding over hours. Characterizing complex vocalizations can require considerable effort and a deep intuition about each species’ vocal behavior. Even with a great deal of experience, human characterizations of animal communication can be affected by human perceptual biases. We present a set of computational methods for projecting animal vocalizations into low dimensional latent representational spaces that are directly learned from the spectrograms of vocal signals. We apply these methods to diverse datasets from over 20 species, including humans, bats, songbirds, mice, cetaceans, and nonhuman primates. Latent projections uncover complex features of data in visually intuitive and quantifiable ways, enabling high-powered comparative analyses of vocal acoustics. We introduce methods for analyzing vocalizations as both discrete sequences and as continuous latent variables. Each method can be used to disentangle complex spectro-temporal structure and observe long-timescale organization in communication. Of the thousands of species that communicate vocally, the repertoires of only a tiny minority have been characterized or studied in detail. This is due, in large part, to traditional analysis methods that require a high level of expertise that is hard to develop and often species-specific. Here, we present a set of unsupervised methods to project animal vocalizations into latent feature spaces to quantitatively compare and develop visual intuitions about animal vocalizations. We demonstrate these methods across a series of analyses over 19 datasets of animal vocalizations from 29 different species, including songbirds, mice, monkeys, humans, and whales. We show how learned latent feature spaces untangle complex spectro-temporal structure, enable cross-species comparisons, and uncover high-level attributes of vocalizations such as stereotypy in vocal element clusters, population regiolects, coarticulation, and individual identity.
DOI: 10.1038/nbt.4314
发表时间: 2019-01-01
影响因子: 46.9
作者:
Becht, Etienne;McInnes, Leland;Newell, Evan W.
通讯作者: Newell, Evan W.
DOI: 10.1163/156853988x00322
发表时间: 1988-12-01
期刊: BEHAVIOUR
影响因子: 1.3
作者:
ADRETHAUSBERGER, M;JENKINS, PF
通讯作者: JENKINS, PF
DOI: 10.1073/pnas.1607601113
发表时间: 2016-10-18
影响因子: 11.1
作者:
Berman, Gordon J.;Bialek, William;Shaevitz, Joshua W.
通讯作者: Shaevitz, Joshua W.
DOI: 10.1016/j.ecoinf.2015.01.007
发表时间: 2015-05-01
影响因子: 5.1
作者:
Arriaga, Julio G.;Cody, Martin L.;Taylor, Charles E.
通讯作者: Taylor, Charles E.
DOI: 10.1073/pnas.1515380113
发表时间: 2016-02-09
影响因子: 11.1
作者:
Bregman, Micah R.;Patel, Aniruddh D.;Gentner, Timothy Q.
通讯作者: Gentner, Timothy Q.