Multisensory Interactions in Speech Perception
Multisensory Interactions in Speech Perception
复制标题
言语感知中的多感官交互
DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
S. Soto
中科院分区:
文献类型:
--
作者:
J. Navarra;H. H. Yeung;J. Werker;S. Soto
The perception of someone talking provides correlated input to more than one sensory modality (mainly vision and audition) simultaneously. There are many everyday situations (such as face-to-face conversations, watching television, or videoconferencing) in which linguistically reliable information can be obtained from the sight of the speaker. The fact that one often communicates effectively in the absence of any visual cue (e.g., talking over the telephone) perhaps leads to the simple inference that these visual speech cues are completely redundant with respect to the concurrent acoustic input or even useless in their linguistic and informational relevance. Nevertheless, empirical evidence accumulated over the last few decades provides solid grounds to dismiss this subjective impression and instead supports the conclusion that there is important information from vision that, when accessible, complements and supplements the acoustic speech signal. Vision carries substantial, and linguistically relevant, cues about the spoken signal. For example, research and clinical/educational practice with deaf individuals have repeatedly demonstrated the benefits of lipreading (or speechreading) under conditions of hearing loss (see Auer, 2010, for a review). Normally hearing individuals also display a remarkable sensitivity to these visual speech cues. For instance, the fact that adults, and even infants as young as 4 months, are capable of discriminating between silent faces articulating sentences in different languages (e.g., English and French) makes us think that the sensitivity to visual speech information arises as a part of normal development and not only as a compensatory strategy to cope with acoustic impairment (Weikum et al., 2007; see also Soto-Faraco et al., 2007; see figure 24.1). Many different linguistic cues (including both segmental and suprasegmental) can, in fact, be retrieved from visual speech articulations (Bernstein, Eberhardt, & Demorest, 1989; Jiang, Auer, Alwan, Keating, & Bernstein, 2007; VatikiotisBateson, Munhall, Kasahara, Garcia, & Yehia,1996; Yehia, Kuratate, & Vatikiotis-Bateson, 2002; and see chapter 23, in this volume, by Vatikiotis-Bateson and Munhall) and from other visible correlates such as head motion (Hadar, Steiner, Grant, & Rose, 1983, 1984; Munhall, Jones, Callan. Kuratate, & Vatikiotis-Bateson, 2004). An important question, however, is whether and how this visual source of information about speech is combined with auditory speech when they are both present. The pioneering work by Sumby and Pollack (1954; see also Cotton, 1935) represents the first successful attempt to assess the role of dynamic facial information on the comprehension of a spoken message. Using a clever setup, these authors demonstrated that the perception of acoustically presented words masked with noise improved substantially when the speaker’s facial movements were available to the perceiver (see also Grant & Greenberg, 2001 and Ross, Saint-Amour, Leavitt, Javitt, & Foxe, 2007, for other demonstrations of visual enhancement of auditory speech perception). This type of result reveals that observers can exploit (and thus benefit from) the informational correspondence between visual and acoustic aspects of the speech signal when needed. Moreover, abundant research suggests that visual speech exerts a substantial impact on speech perception even under good acoustic conditions (e.g., Reisberg, McLean, & Goldfield, 1987), that is, not only when the acoustical signal is degraded. McGurk and MacDonald (1976) demonstrated this point in a highly influential study, showing that a syllable that is heard as /ba/ when presented in isolation is often heard as /da/ when dubbed onto a video clip of a face silently articulating the syllable [ga] (see figure 24.2). This added value of audiovisual (AV) integration has also been demonstrated in more subtle ways, without the need for artificially induced intersensory conflict. For example, the combination of visual and auditory speech can make us more sensitive to nonnative phonemic distinctions that are difficult to discern on the basis of just visual or auditory information alone (Navarra & Soto-Faraco, 2007; see also Teinonen, Aslin, Alku, & Csibra, 2008). Developmental research has shown that these AV speech perception abilities are
登录
查看更多内容
影响因子:
2.4
作者:
Stevens, KN
通讯作者:
Stevens, KN
DOI:
10.1037//0096-1523.17.3.816
发表时间:
1991
期刊:
Journal of experimental psychology. Human perception and performance
影响因子:
--
作者:
Fowler,CA;Dekle,DJ
通讯作者:
Dekle,DJ
影响因子:
3.7
作者:
T. Wright;K. Pelphrey;T. Allison;M. McKeown;G. McCarthy
通讯作者:
T. Wright;K. Pelphrey;T. Allison;M. McKeown;G. McCarthy
DOI:
10.1121/1.423909
发表时间:
1998
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
作者:
AuerJr,ET;Bernstein,LE;Coulter,DC
通讯作者:
Coulter,DC
DOI:
10.1037//0096-1523.21.6.1409
发表时间:
1995
期刊:
Journal of experimental psychology. Human perception and performance
影响因子:
--
作者:
Green,KP;Gerdeman,A
通讯作者:
Gerdeman,A