Neural network vowel-recognition jointly using voice features and mouth shape image
Neural network vowel-recognition jointly using voice features and mouth shape image
复制标题
联合使用语音特征和嘴形图像的神经网络元音识别
DOI:
10.1016/0031-3203(91)90089-n
复制
发表时间:
1991
期刊:
影响因子:
--
通讯作者:
K. Okazaki
中科院分区:
文献类型:
--
作者:
Jian;S. Tamura;H. Mitsumoto;H. Kawai;K. Kurosu;K. Okazaki
This paper describes a neural approach intended to improve the performance of an automatic speech recognition system for unrestricted speakers by using not only voice sound features but also image features of the mouth shape. In particular, we used the natural sample voice signals and mouth shape images that were acquired in the general environment, neither in the sound isolation room nor under specific lighting conditions. The FFT power spectrum of acoustic speech was used as the voice feature. In addition, the gray level image, binary image and geometrical shape features of the mouth were used as the compensatory information, and compared which kinds of image features are effective. For unrestricted speakers, a vowel recognition rate of about 80% was obtained using only voice features, but this increased to some 92% when voice features plus binary images were used. This method can be applied not only to the improvement of voice recognition, but also to aid the communication of hearing-impaired people.