Multisensory Interactions in Speech Perception

Multisensory Interactions in Speech Perception
复制标题

言语感知中的多感官交互

DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
S. Soto
S. Soto
中科院分区:
--
文献类型:
--
作者:
J. Navarra;H. H. Yeung;J. Werker;S. Soto

文献摘要

参考文献

被引文献

相似文献

对某人说话的感知同时为不止一种感觉通道(主要是视觉和听觉)提供了相关的输入。在许多日常情况下(如面对面的谈话、看电视或视频会议),可以从说话者的视线中获得语言上可靠的信息。一个人经常在没有任何视觉提示(例如,通过电话交谈)的情况下进行有效沟通的事实可能导致一个简单的推论,即这些视觉语音提示相对于同时的声学输入是完全多余的,甚至在它们的语言和信息相关性方面是无用的。然而,过去几十年积累的经验证据提供了坚实的理由来驳斥这种主观印象,相反,它支持这样的结论,即来自视觉的重要信息在可获得时,补充和补充了声学语音信号。视觉携带着大量的、与语言相关的关于口头信号的线索。例如,对聋人的研究和临床/教育实践一再证明,在听力丧失的情况下唇读(或朗读演讲稿)的好处(见Auer,2010,综述)。正常情况下,听力正常的人对这些视觉语音提示也表现出非凡的敏感性。例如,成年人,甚至4个月大的婴儿,都能够区分用不同语言(如英语和法语)发出句子的无声面孔,这一事实使我们认为,对视觉言语信息的敏感性是正常发育的一部分,而不仅仅是应对听觉障碍的一种补偿策略(Wekum等人,2007年;另见Soto-Faraco等人,2007;见图24.1)。事实上,许多不同的语言线索(包括音段和超音段)可以从视觉语言表达中检索到(Bernstein,Eberhardt&Demorest,1989;酱,Auer,Alwan,Kating,&Bernstein,2007;Vatikiotis Bateson,MunHall,Kasahara,Garcia,&Yehia,1996;Yehia,Kuratate,&Vatikiotis-Bateson,2002;以及见本卷第23章,Vatikiotis-Bateson和MunHall,Vatikiotis-Bateson和MunHall)以及其他可视关联,如头部运动(Hadar,Steiner,Grant,&Rose,1983;MunHall,Jones,Callan,1984)。Kuratate和Vatikiotis-Bateson,2004)。然而,一个重要的问题是,当语音的视觉信息源和听觉语音都存在时,它们是否以及如何结合在一起。Sumby和Pollack的开创性工作(1954;另见Cotton,1935)代表了评估动态面部信息在理解口语信息中的作用的第一次成功尝试。使用一种巧妙的设置,这些作者证明了当说话者的面部动作对接受者可用时,通过声音呈现的被噪声掩盖的单词的感知显著改善(另见Grant&Greenberg,2001和Ross,Saint-Amour,Leavitt,Javitt,&Foxe,2007,关于听觉言语感知的视觉增强的其他演示)。这种类型的结果表明,观察者可以在需要时利用语音信号的视觉和听觉方面之间的信息对应关系(并因此从中受益)。此外,大量研究表明,即使在良好的声学条件下(例如Reisberg,McLean,&Goldfield,1987),视觉言语也会对言语知觉产生重大影响,也就是说,不仅是在声音信号退化的情况下。McGurk和MacDonald(1976)在一项非常有影响力的研究中证明了这一点,该研究表明,当孤立地呈现一个音节时,被听到的音节通常被听到为/da/,当被配音到一个面部默默地发出音节[ga]的视频剪辑上时,通常被听到为/da/(见图24.2)。视听(AV)整合的这种附加值也以更微妙的方式得到了证明,而不需要人为地诱导感官间冲突。例如,视觉和听觉的结合可以使我们对仅凭视觉或听觉信息很难区分的非母语音素区别更加敏感(Navarra&Soto-Faraco,2007;另见Teinonen,Aslin,ALKU和Csibra,2008)。发展研究表明,这些视听语音感知能力是
The perception of someone talking provides correlated input to more than one sensory modality (mainly vision and audition) simultaneously. There are many everyday situations (such as face-to-face conversations, watching television, or videoconferencing) in which linguistically reliable information can be obtained from the sight of the speaker. The fact that one often communicates effectively in the absence of any visual cue (e.g., talking over the telephone) perhaps leads to the simple inference that these visual speech cues are completely redundant with respect to the concurrent acoustic input or even useless in their linguistic and informational relevance. Nevertheless, empirical evidence accumulated over the last few decades provides solid grounds to dismiss this subjective impression and instead supports the conclusion that there is important information from vision that, when accessible, complements and supplements the acoustic speech signal. Vision carries substantial, and linguistically relevant, cues about the spoken signal. For example, research and clinical/educational practice with deaf individuals have repeatedly demonstrated the benefits of lipreading (or speechreading) under conditions of hearing loss (see Auer, 2010, for a review). Normally hearing individuals also display a remarkable sensitivity to these visual speech cues. For instance, the fact that adults, and even infants as young as 4 months, are capable of discriminating between silent faces articulating sentences in different languages (e.g., English and French) makes us think that the sensitivity to visual speech information arises as a part of normal development and not only as a compensatory strategy to cope with acoustic impairment (Weikum et al., 2007; see also Soto-Faraco et al., 2007; see figure 24.1). Many different linguistic cues (including both segmental and suprasegmental) can, in fact, be retrieved from visual speech articulations (Bernstein, Eberhardt, & Demorest, 1989; Jiang, Auer, Alwan, Keating, & Bernstein, 2007; VatikiotisBateson, Munhall, Kasahara, Garcia, & Yehia,1996; Yehia, Kuratate, & Vatikiotis-Bateson, 2002; and see chapter 23, in this volume, by Vatikiotis-Bateson and Munhall) and from other visible correlates such as head motion (Hadar, Steiner, Grant, & Rose, 1983, 1984; Munhall, Jones, Callan. Kuratate, & Vatikiotis-Bateson, 2004). An important question, however, is whether and how this visual source of information about speech is combined with auditory speech when they are both present. The pioneering work by Sumby and Pollack (1954; see also Cotton, 1935) represents the first successful attempt to assess the role of dynamic facial information on the comprehension of a spoken message. Using a clever setup, these authors demonstrated that the perception of acoustically presented words masked with noise improved substantially when the speaker’s facial movements were available to the perceiver (see also Grant & Greenberg, 2001 and Ross, Saint-Amour, Leavitt, Javitt, & Foxe, 2007, for other demonstrations of visual enhancement of auditory speech perception). This type of result reveals that observers can exploit (and thus benefit from) the informational correspondence between visual and acoustic aspects of the speech signal when needed. Moreover, abundant research suggests that visual speech exerts a substantial impact on speech perception even under good acoustic conditions (e.g., Reisberg, McLean, & Goldfield, 1987), that is, not only when the acoustical signal is degraded. McGurk and MacDonald (1976) demonstrated this point in a highly influential study, showing that a syllable that is heard as /ba/ when presented in isolation is often heard as /da/ when dubbed onto a video clip of a face silently articulating the syllable [ga] (see figure 24.2). This added value of audiovisual (AV) integration has also been demonstrated in more subtle ways, without the need for artificially induced intersensory conflict. For example, the combination of visual and auditory speech can make us more sensitive to nonnative phonemic distinctions that are difficult to discern on the basis of just visual or auditory information alone (Navarra & Soto-Faraco, 2007; see also Teinonen, Aslin, Alku, & Csibra, 2008). Developmental research has shown that these AV speech perception abilities are
DOI: 10.1121/1.1458026
发表时间: 2002-04-01
影响因子: 2.4
作者:
Stevens, KN
通讯作者: Stevens, KN
用眼睛和手听:对语音感知的跨模式贡献。
DOI: 10.1037//0096-1523.17.3.816
发表时间: 1991
期刊: Journal of experimental psychology. Human perception and performance
影响因子: --
作者:
Fowler,CA;Dekle,DJ
通讯作者: Dekle,DJ
DOI: 10.1093/cercor/13.10.1034
发表时间: 2003-10
期刊: Cerebral cortex
影响因子: 3.7
作者:
T. Wright;K. Pelphrey;T. Allison;M. McKeown;G. McCarthy
通讯作者: T. Wright;K. Pelphrey;T. Allison;M. McKeown;G. McCarthy
语音基频的时间和时空振动触觉显示:对正常听力和听力受损个体的新型振动触觉语音感知辅助的初步评估。
DOI: 10.1121/1.423909
发表时间: 1998
期刊: The Journal of the Acoustical Society of America
影响因子: --
作者:
AuerJr,ET;Bernstein,LE;Coulter,DC
通讯作者: Coulter,DC
DOI: 10.1111/j.1467-8624.1992.tb01661.x
发表时间: 1992-08
期刊: Child development
影响因子: 4.6
作者:
Nelson H. Soken;A. Pick
通讯作者: Nelson H. Soken;A. Pick