Bidirectional Generative Adversarial Representation Learning for Natural Stimulus Synthesis

Bidirectional Generative Adversarial Representation Learning for Natural Stimulus Synthesis
复制标题

DOI:
10.1101/2023.10.17.562789
复制
发表时间:
2023-10
期刊:
bioRxiv
影响因子:
--
通讯作者:
Johnny Reilly;John D. Goodwin;Sihao Lu;Andriy S. Kozlov
Johnny Reilly;John D. Goodwin;Sihao Lu;Andriy S. Kozlov
中科院分区:
其他
文献类型:
--
作者:
Johnny Reilly;John D. Goodwin;Sihao Lu;Andriy S. Kozlov

文献摘要

相似文献

数以千计的物种使用声音信号相互交流。发声携带着丰富的信息,但表征和分析这些复杂的高维信号是困难的,而且容易受到人类偏见的影响。此外,动物发声是行为学上相关的刺激,其听觉神经元的表征是感觉神经科学的重要研究课题。一种可以有效地产生自然发声波形的方法将提供无限的刺激,以探索神经元计算。虽然无监督学习方法允许将发声投射到从波形本身学习的低维潜在空间中,并且生成建模允许合成用于下游任务的新颖发声,但是目前没有将联合收割机这些任务组合以产生用于刺激回放的自然发声波形的方法。在本文中,我们展示了BiWaveGAN:一种双向生成对抗网络(GAN),能够从小鼠身上学习超声发声(USVs)的潜在表示。我们表明,BiWaveGAN可以用来生成,并插值,现实的发声波形。然后,我们使用这些合成的刺激沿着与自然USV探测小鼠听觉皮层神经元的感觉输入空间。我们表明,从我们的方法产生的刺激引起神经元的反应,有效地作为真实的发声,并产生具有相同的预测能力的感受野。BiWaveGAN不限于小鼠USV,但可用于合成任何动物物种的自然发声,并在相同或不同物种的发声之间进行插值,这可能有助于探索行为学相关听觉信号表示中的分类边界。
Thousands of species use vocal signals to communicate with one another. Vocalisations carry rich information, yet characterising and analysing these complex, high-dimensional signals is difficult and prone to human bias. Moreover, animal vocalisations are ethologically relevant stimuli whose representation by auditory neurons is an important subject of research in sensory neuroscience. A method that can efficiently generate naturalistic vocalisation waveforms would offer an unlimited supply of stimuli with which to probe neuronal computations. While unsupervised learning methods allow for the projection of vocalisations into low-dimensional latent spaces learned from the waveforms themselves, and generative modelling allows for the synthesis of novel vocalisations for use in downstream tasks, there is currently no method that would combine these tasks to produce naturalistic vocalisation waveforms for stimulus playback. In this paper, we demonstrate BiWaveGAN: a bidirectional Generative Adversarial Network (GAN) capable of learning a latent representation of ultrasonic vocalisations (USVs) from mice. We show that BiWaveGAN can be used to generate, and interpolate between, realistic vocalisation waveforms. We then use these synthesised stimuli along with natural USVs to probe the sensory input space of mouse auditory cortical neurons. We show that stimuli generated from our method evoke neuronal responses as effectively as real vocalisations, and produce receptive fields with the same predictive power. BiWaveGAN is not restricted to mouse USVs but can be used to synthesise naturalistic vocalisations of any animal species and interpolate between vocalisations of the same or different species, which could be useful for probing categorical boundaries in representations of ethologically relevant auditory signals.