Reading speech from still and moving faces: The neural substrates of visible speech

Reading speech from still and moving faces: The neural substrates of visible speech
复制标题

DOI:
10.1162/089892903321107828
复制
发表时间:
2003-01-01
影响因子:
3.2
通讯作者:
Campbell, R
Campbell, R
中科院分区:
医学3区
文献类型:
--
作者:
Calvert, GA;Campbell, R

文献摘要

被引文献

相似文献

言语既可以用耳朵也可以用眼睛感知。与听到的语音不同,一些看到的语音手势可以在静止图像序列中捕获。先前的研究表明,在听力正常的人中,自然时变的无声的可见语音可以进入听觉皮层(左上级颞区)。本研究利用功能性磁共振成像(fMRI)技术,探索了当视觉语言被剥夺其时变特征时,该回路被激活的程度。在扫描仪中,听力参与者被指示在其他单音节词中寻找预先指定的视觉语言目标序列(“voo”或“ahv”)。在一种情况下,图像序列包括示出心尖姿势的一系列静止关键帧(例如,用于“v”和“oo”[来自目标]或“ee”和“m”[即,从非目标音节])。在另一种情况下,自然的语音运动相同的整体段的持续时间被seeed.In对比的基线条件,其中字母“V”被叠加在一个休息的脸,静止的语音人脸图像产生激活后皮层区域与感知的生物运动,尽管缺乏明显的运动在语音图像序列。在传统的语音处理区域也检测到激活,包括左下额叶(布罗卡)区,左上级颞沟(STS),和左缘上回(韦尼克区的背面)。静止的语音序列也产生激活腹侧前运动皮层和前下顶叶沟bilateral.Moving脸产生显着更大的皮层激活比静止的脸序列,在类似的地区。然而,也观察到了静止和移动语音之间的一些差异。在视觉皮层中,静止面孔在初级视觉区(V1/V2)产生相对更多的激活,而视觉运动区(V5/MT+)在更大程度上被运动面孔激活。皮质区激活更多的自然移动说话的面孔包括听觉皮层(布罗德曼的地区41/42;侧部Heschl的gyrus)和左侧STS和额下回。看到语音与正常的时变特性似乎有优先访问“纯粹”的听觉处理区域专门的语言,可能通过收购动态视听整合机制STS。当看到的语音缺乏自然的时变特征时,左颞叶的语音处理系统可能主要通过基于动作的语音表征来实现,在腹侧前运动皮层中实现。
Speech is perceived both by ear and by eye. Unlike heard speech, some seen speech gestures can be captured in stilled image sequences. Previous studies have shown that in hearing people, natural time-varying silent seen speech can access the auditory cortex (left superior temporal regions). Using functional magnetic resonance imaging (fMRI), the present study explored the extent to which this circuitry was activated when seen speech was deprived of its time-varying characteristics.In the scanner, hearing participants were instructed to look for a prespecified visible speech target sequence ("voo" or "ahv") among other monosyllables. In one condition, the image sequence comprised a series of stilled key frames showing apical gestures (e.g., separate frames for "v" and "oo" [from the target] or "ee" and "m" [i.e., from nontarget syllables]). In the other condition, natural speech movement of the same overall segment duration was seen.In contrast to a baseline condition in which the letter "V" was superimposed on a resting face, stilled speech face images generated activation in posterior cortical regions associated with the perception of biological movement, despite the lack of apparent movement in the speech image sequence. Activation was also detected in traditional speech-processing regions including the left inferior frontal (Broca's) area, left superior temporal sulcus (STS), and left supramarginal gyrus (the dorsal aspect of Wernicke's area). Stilled speech sequences also generated activation in the ventral premotor cortex and anterior inferior parietal sulcus bilaterally.Moving faces generated significantly greater cortical activation than stilled face sequences, and in similar regions. However, a number of differences between stilled and moving speech were also observed. in the visual cortex, stilled faces generated relatively more activation in primary visual regions (V1/V2), while visual movement areas (V5/MT+) were activated to a greater extent by moving faces. Cortical regions activated more by naturally moving speaking faces included the auditory cortex (Brodmann's Areas 41/42; lateral parts of Heschl's gyrus) and the left STS and inferior frontal gyrus.Seen speech with normal time-varying characteristics appears to have preferential access to "purely" auditory processing regions specialized for language, possibly via acquired dynamic audiovisual integration mechanisms in STS. When seen speech lacks natural time-varying characteristics, access to speech-processing systems in the left temporal lobe may be achieved predominantly via action-based speech representations, realized in the ventral premotor cortex.