Neural network vowel-recognition jointly using voice features and mouth shape image

Neural network vowel-recognition jointly using voice features and mouth shape image
复制标题

联合使用语音特征和嘴形图像的神经网络元音识别

DOI:
10.1016/0031-3203(91)90089-n
复制
发表时间:
1991
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
K. Okazaki
K. Okazaki
中科院分区:
--
文献类型:
--
作者:
Jian;S. Tamura;H. Mitsumoto;H. Kawai;K. Kurosu;K. Okazaki

文献摘要

被引文献

相似文献

本文介绍了一种神经网络方法,该方法不仅利用语音特征,而且利用嘴巴形状的图像特征来提高自动语音识别系统的性能。特别地,我们使用了在一般环境下获得的自然样本语音信号和嘴形图像,而不是在隔声室或特定照明条件下。利用声学语音的FFT功率谱作为语音特征。此外,利用灰度图像、二值图像和嘴的几何形状特征作为补偿信息,比较了几种图像特征的有效性。对于不受限制的说话者,仅使用语音特征获得的元音识别率约为80%,但当使用语音特征加二值图像时,这一比例增加到92%左右。该方法不仅可以用于语音识别的改进,还可以帮助听障人士进行交流。
This paper describes a neural approach intended to improve the performance of an automatic speech recognition system for unrestricted speakers by using not only voice sound features but also image features of the mouth shape. In particular, we used the natural sample voice signals and mouth shape images that were acquired in the general environment, neither in the sound isolation room nor under specific lighting conditions. The FFT power spectrum of acoustic speech was used as the voice feature. In addition, the gray level image, binary image and geometrical shape features of the mouth were used as the compensatory information, and compared which kinds of image features are effective. For unrestricted speakers, a vowel recognition rate of about 80% was obtained using only voice features, but this increased to some 92% when voice features plus binary images were used. This method can be applied not only to the improvement of voice recognition, but also to aid the communication of hearing-impaired people.