EARSHOT: A minimal network model of human speech recognition that operates on real speech

EARSHOT: A minimal network model of human speech recognition that operates on real speech
复制标题

DOI:
--
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
J. Magnuson;Heejo You;J. Rueckl;Paul D. Allopenna;Monica Li;Sahil Luthra;Rachael Steiner;Hosung Nam;M. Escabí;K. Brown;Rachel M. Theodore;Nicholas Monto
J. Magnuson;Heejo You;J. Rueckl;Paul D. Allopenna;Monica Li;Sahil Luthra;Rachael Steiner;Hosung Nam;M. Escabí;K. Brown;Rachel M. Theodore;Nicholas Monto
中科院分区:
其他
文献类型:
--
作者:
J. Magnuson;Heejo You;J. Rueckl;Paul D. Allopenna;Monica Li;Sahil Luthra;Rachael Steiner;Hosung Nam;M. Escabí;K. Brown;Rachel M. Theodore;Nicholas Monto

文献摘要

相似文献

尽管没有不变性问题(声学和知觉之间的多对多映射),但我们经历了语音恒定,并通常感知说话人的意图。人类语音识别的模型避开了这个问题,使用抽象、理想化的输入,推迟了处理真实语音的挑战。相比之下,由深度学习网络支持的自动语音识别实现了强大的、真实世界的语音识别。然而,深度学习架构和训练方案的复杂性使得很难使用它们来提供对可能支持人类语音识别的机制的直接见解。我们开发了一个简单的网络,它借用了自动语音识别的一个元素(长短期记忆节点,它为短跨度和长跨度提供动态记忆)。这使得网络能够学习以高精度将来自多个说话者的真实语音映射到语义目标。在人类上颞回中出现了类似于语音组织反应的内部表征,这表明尽管没有对语音或音素目标进行明确的训练,但该模型发展了一种分布式语音编码。处理真实语音的能力是人类语音识别认知模型的一大进步。
Despite the lack of invariance problem (the many-to-many mapping between acoustics and percepts), we experience phonetic constancy and typically perceive what a speaker intends. Models of human speech recognition have sidestepped this problem, working with abstract, idealized inputs and deferring the challenge of working with real speech. In contrast, automatic speech recognition powered by deep learning networks have allowed robust, real-world speech recognition. However, the complexities of deep learning architectures and training regimens make it difficult to use them to provide direct insights into mechanisms that may support human speech recognition. We developed a simple network that borrows one element from automatic speech recognition (long short-term memory nodes, which provide dynamic memory for short and long spans). This allows the network to learn to map real speech from multiple talkers to semantic targets with high accuracy. Internal representations emerge that resemble phonetically-organized responses in human superior temporal gyrus, suggesting that the model develops a distributed phonological code despite no explicit training on phonetic or phonemic targets. The ability to work with real speech is a major advance for cognitive models of human speech recognition.