Robust speech perception: recognize the familiar, generalize to the similar, and adapt to the novel.

Robust speech perception: recognize the familiar, generalize to the similar, and adapt to the novel.
复制标题

DOI:
10.1037/a0038695
复制
发表时间:
2015-04
影响因子:
5.4
通讯作者:
Jaeger TF
Jaeger TF
中科院分区:
心理学1区
文献类型:
--
作者:
Kleinschmidt DF;Jaeger TF

文献摘要

被引文献

相似文献

成功的语音感知需要听者将声音信号映射到语言类别。这些映射不仅是概率性的,而且根据情况而变化。例如,一个说话者的/p/可能在物理上与另一个说话者的/b/无法区分(参见缺乏不变性)。我们描述了这种主观非平稳世界所带来的计算问题,并提出语音感知系统通过(1)识别以前遇到的情况,(2)根据以前的类似经验推广到其他情况,以及(3)适应新情况来克服这一挑战。我们在理想的适配器框架中形式化了这一建议:(1)到(3)可以理解为在不确定的情况下对当前说话者的适当生成模型的推断,从而在缺乏不变性的情况下促进鲁棒语音感知。我们关注理想适配器的两个关键方面。首先,在明显偏离以往经验的情况下,听众需要适应。我们开发了一种增量适应的分布式(信念更新)学习模型。该模型对已知的和新的语音适应数据提供了很好的拟合,包括感知重新校准和选择性适应。其次,稳健的语音识别要求听者学会表征语音信号中跨情境可变性的结构化成分。我们将讨论理想适配器的这两个方面如何为跨说话者和说话者群体(例如,口音和方言)的适应、说话者特异性和泛化提供统一的解释。理想的适配器为未来研究语音感知和适应以及更广泛的语言理解提供了一个指导框架。
Successful speech perception requires that listeners map the acoustic signal to linguistic categories. These mappings are not only probabilistic, but change depending on the situation. For example, one talker’s /p/ might be physically indistinguishable from another talker’s /b/ (cf. lack of invariance). We characterize the computational problem posed by such a subjectively non-stationary world and propose that the speech perception system overcomes this challenge by (1) recognizing previously encountered situations, (2) generalizing to other situations based on previous similar experience, and (3) adapting to novel situations. We formalize this proposal in the ideal adapter framework: (1) to (3) can be understood as inference under uncertainty about the appropriate generative model for the current talker, thereby facilitating robust speech perception despite the lack of invariance. We focus on two critical aspects of the ideal adapter. First, in situations that clearly deviate from previous experience, listeners need to adapt. We develop a distributional (belief-updating) learning model of incremental adaptation. The model provides a good fit against known and novel phonetic adaptation data, including perceptual recalibration and selective adaptation. Second, robust speech recognition requires listeners learn to represent the structured component of cross-situation variability in the speech signal. We discuss how these two aspects of the ideal adapter provide a unifying explanation for adaptation, talker-specificity, and generalization across talkers and groups of talkers (e.g., accents and dialects). The ideal adapter provides a guiding framework for future investigations into speech perception and adaptation, and more broadly language comprehension.