Automatic audiovisual integration in speech perception

Automatic audiovisual integration in speech perception
复制标题

DOI:
10.1007/s00221-005-0008-z
复制
发表时间:
2005-11-01
影响因子:
2
通讯作者:
Cattaneo, L
Cattaneo, L
中科院分区:
医学4区
文献类型:
--
作者:
Gentilucci, M;Cattaneo, L

文献摘要

被引文献

相似文献

两个实验的目的是确定视觉和听觉输入的功能是否总是合并到语音的感知表示,以及这种视听集成是否基于跨模态绑定功能或模仿。在麦格克范式中,观察者被要求大声重复一个演员发出的一串音素(音素串的声学呈现),而演员的嘴则模仿不同字符串的发音(视觉呈现)。在对照实验中,参与者阅读相同的打印字母串。该条件旨在分析语音模式和嘴唇运动学控制模仿。在控制实验和一致的视听演示,即当发音嘴的姿态是一致的发射的字符串的电话,语音频谱和嘴唇的运动学变化根据发音字符串的音素。在McGurk范式中,参与者没有意识到视觉和听觉刺激之间的不一致。对参与者的口语反应的声学分析显示了三种不同的模式:两种刺激的融合(麦格克效应),声学上呈现的音素串的重复,以及与演员模仿的嘴部手势相对应的音素串的重复。然而,对后两种反应的分析表明,参与者的语音频谱的共振峰2总是不同于在一致的视听呈现中记录的值。它接近在另一模态中呈现的音素串的共振峰2的值,这显然被忽略了。嘴唇运动学的参与者重复的字符串的音素声学呈现的影响,由演员模仿的嘴唇运动的观察,但只有当发音唇辅音。数据进行了讨论,有利于假设的视觉和声学输入的功能总是有助于表示的一串音素和跨模态集成发生提取口发音特有的字符串音素的发音功能。
Two experiments aimed to determine whether features of both the visual and acoustical inputs are always merged into the perceived representation of speech and whether this audiovisual integration is based on either cross-modal binding functions or on imitation. In a McGurk paradigm, observers were required to repeat aloud a string of phonemes uttered by an actor (acoustical presentation of phonemic string) whose mouth, in contrast, mimicked pronunciation of a different string (visual presentation). In a control experiment participants read the same printed strings of letters. This condition aimed to analyze the pattern of voice and the lip kinematics controlling for imitation. In the control experiment and in the congruent audiovisual presentation, i.e. when the articulation mouth gestures were congruent with the emission of the string of phones, the voice spectrum and the lip kinematics varied according to the pronounced strings of phonemes. In the McGurk paradigm the participants were unaware of the incongruence between visual and acoustical stimuli. The acoustical analysis of the participants' spoken responses showed three distinct patterns: the fusion of the two stimuli (the McGurk effect), repetition of the acoustically presented string of phonemes, and, less frequently, of the string of phonemes corresponding to the mouth gestures mimicked by the actor. However, the analysis of the latter two responses showed that the formant 2 of the participants' voice spectra always differed from the value recorded in the congruent audiovisual presentation. It approached the value of the formant 2 of the string of phonemes presented in the other modality, which was apparently ignored. The lip kinematics of the participants repeating the string of phonemes acoustically presented were influenced by the observation of the lip movements mimicked by the actor, but only when pronouncing a labial consonant. The data are discussed in favor of the hypothesis that features of both the visual and acoustical inputs always contribute to the representation of a string of phonemes and that cross-modal integration occurs by extracting mouth articulation features peculiar for the pronunciation of that string of phonemes.