Development of a silent speech interface driven by ultrasound and optical images of the tongue and lips

Development of a silent speech interface driven by ultrasound and optical images of the tongue and lips
复制标题

DOI:
10.1016/j.specom.2009.11.004
复制
发表时间:
2010-04-01
影响因子:
3.2
通讯作者:
Stone, Maureen
Stone, Maureen
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hueber, Thomas;Benaroya, Elie-Laurent;Stone, Maureen

文献摘要

被引文献

相似文献

本文提出了一种分段声码器驱动的超声和光学图像(标准CCD摄像机)的舌头和嘴唇的“无声的语音接口”的应用程序,可用于喉切除术的患者或无声的通信。该系统是建立在一个视听词典相关联的视觉声学观察每个语音类。视觉特征提取从超声图像的舌头和视频图像的嘴唇使用PCA为基础的图像编码技术。每个语音类的视觉观察是由连续的Hynth建模。然后,该系统将电话识别阶段与基于语料库的合成相结合。在识别阶段,视觉障碍被用来识别语音目标的视觉特征序列。在合成阶段,这些语音目标约束字典搜索的序列的双音素,最大限度地提高相似性的输入测试数据在视觉空间中,在声学域中的级联成本。从训练语料中提取韵律模板,并使用“谐波加噪声模型”拼接合成技术生成最终的语音波形。实验结果是基于一个视听数据库,其中包含I h的连续语音从两个扬声器。(C)2009 Elsevier B.V.保留所有权利。
This article presents a segmental vocoder driven by ultrasound and optical images (standard CCD camera) of the tongue and lips for a "silent speech interface" application, usable either by a laryngectomized patient or for silent communication. The system is built around an audio visual dictionary which associates visual to acoustic observations for each phonetic class. Visual features are extracted from ultrasound images of the tongue and from video images of the lips using a PCA-based image coding technique. Visual observations of each phonetic class are modeled by continuous HMMs. The system then combines a phone recognition stage with corpus-based synthesis. In the recognition stage, the visual HMMs are used to identify phonetic targets in a sequence of visual features. In the synthesis stage, these phonetic targets constrain the dictionary search for the sequence of diphones that maximizes similarity to the input test data in the visual space, subject to a concatenation cost in the acoustic domain. A prosody-template is extracted from the training corpus, and the final speech waveform is generated using "Harmonic plus Noise Model" concatenative synthesis techniques. Experimental results are based on an audiovisual database containing I h of continuous speech from each of two speakers. (C) 2009 Elsevier B.V. All rights reserved.