EARSHOT: A Minimal Neural Network Model of Incremental Human Speech Recognition

EARSHOT: A Minimal Neural Network Model of Incremental Human Speech Recognition
复制标题

DOI:
10.1111/cogs.12823
复制
发表时间:
2020-04-01
期刊:
影响因子:
2.5
通讯作者:
Rueckl, Jay G.
Rueckl, Jay G.
中科院分区:
心理学3区
文献类型:
--
作者:
Magnuson, James S.;You, Heejo;Rueckl, Jay G.

文献摘要

被引文献

相似文献

尽管没有不变性问题(声学和知觉之间的多对多映射),但人类听者体验到语音的恒定,并通常感知说话人的意图。大多数人类语音识别(HSR)模型都回避了这个问题,使用抽象的、理想化的输入,推迟了处理真实语音的挑战。相比之下,精心设计的深度学习网络可以实现强大的、真实世界的自动语音识别(ASR)。然而,深度学习架构和培训方案的复杂性使得很难使用它们来提供对可能支持高铁的机制的直接见解。在这篇简短的文章中,我们报告了一个两层网络的初步结果,该网络借用了ASR的一个元素,即长期短期记忆节点,它为一系列时间跨度提供动态记忆。这使得该模型能够学习以高精度将来自多个说话者的真实语音映射到语义目标,以及类似于人类的词汇获取和语音竞争的时间进程。在人类上颞回中出现了类似于语音组织反应的内部表征,这表明尽管没有对语音或音素目标进行明确的训练,但该模型发展了一种分布式语音代码。处理真实语音的能力是高铁认知模型的一大进步。
Despite the lack of invariance problem (the many-to-many mapping between acoustics and percepts), human listeners experience phonetic constancy and typically perceive what a speaker intends. Most models of human speech recognition (HSR) have side-stepped this problem, working with abstract, idealized inputs and deferring the challenge of working with real speech. In contrast, carefully engineered deep learning networks allow robust, real-world automatic speech recognition (ASR). However, the complexities of deep learning architectures and training regimens make it difficult to use them to provide direct insights into mechanisms that may support HSR. In this brief article, we report preliminary results from a two-layer network that borrows one element from ASR, long short-term memory nodes, which provide dynamic memory for a range of temporal spans. This allows the model to learn to map real speech from multiple talkers to semantic targets with high accuracy, with human-like timecourse of lexical access and phonological competition. Internal representations emerge that resemble phonetically organized responses in human superior temporal gyrus, suggesting that the model develops a distributed phonological code despite no explicit training on phonetic or phonemic targets. The ability to work with real speech is a major advance for cognitive models of HSR.