Speech perception at the interface of neurobiology and linguistics

Speech perception at the interface of neurobiology and linguistics
复制标题

DOI:
10.1098/rstb.2007.2160
复制
发表时间:
2008-03-12
影响因子:
6.3
通讯作者:
van Wassenhove, Virginie
van Wassenhove, Virginie
中科院分区:
生物学1区
文献类型:
--
作者:
Poeppel, David;Idsardi, William J.;van Wassenhove, Virginie

文献摘要

被引文献

相似文献

语音感知由一组计算组成,这些计算将连续变化的声波波形作为输入,并产生离散的表征,这些表征与存储在长期记忆中的词汇表征联系起来作为输出。由于语音感知所识别的感知对象进入后续的语言计算,用于词汇表征和处理的格式从根本上制约了语音感知过程。因此,在某种程度上,言语感知理论必须与词汇表征理论紧密相连。在最低限度上,语音感知必须产生与存储的词项平稳而快速地交互的表示。采用MARR的观点,我们论证并为下一步的研究计划提供神经生物学和心理物理学证据。首先,在实现层面上,语音感知是一个多时间分辨率的过程,感知分析在至少两个时间尺度上同时进行(大约。20-80毫秒,大约150-300ms),分别与(亚)音段分析和音节分析相称。其次,在算法层面上,我们认为感知是基于内部前瞻模型进行的,或者使用一种“综合分析”的方法。第三,在计算层面(在Marr的意义上),我们采用的词汇表征理论主要是由音系学研究提供的,并假设词汇在心理词典中是以由不同特征组成的离散片段序列来表征的。该研究计划的一个重要目标是在假定的神经生物学原语(例如时间原语)和那些来自语言研究的原语之间建立联系的假设,最终得出一个在生物上敏感的和在理论上令人满意的言语表征和计算模式。
Speech perception consists of a set of computations that take continuously varying acoustic waveforms as input and generate discrete representations that make contact with the lexical representations stored in long-term memory as output. Because the perceptual objects that are recognized by the speech perception enter into subsequent linguistic computation, the format that is used for lexical representation and processing fundamentally constrains the speech perceptual processes. Consequently, theories of speech perception must, at some level, be tightly linked to theories of lexical representation. Minimally, speech perception must yield representations that smoothly and rapidly interface with stored lexical items. Adopting the perspective of Marr, we argue and provide neurobiological and psychophysical evidence for the following research programme. First, at the implementational level, speech perception is a multi-time resolution process, with perceptual analyses occurring concurrently on at least two time scales (approx. 20 - 80 ms, approx. 150 - 300 ms), commensurate with (sub) segmental and syllabic analyses, respectively. Second, at the algorithmic level, we suggest that perception proceeds on the basis of internal forward models, or uses an 'analysis-by-synthesis' approach. Third, at the computational level (in the sense of Marr), the theory of lexical representation that we adopt is principally informed by phonological research and assumes that words are represented in the mental lexicon in terms of sequences of discrete segments composed of distinctive features. One important goal of the research programme is to develop linking hypotheses between putative neurobiological primitives (e.g. temporal primitives) and those primitives derived from linguistic inquiry, to arrive ultimately at a biologically sensible and theoretically satisfying model of representation and computation in speech.