Audio-Visual Speech Recognition Using Convolutive Bottleneck Networks for a Person with Severe Hearing Loss

Audio-Visual Speech Recognition Using Convolutive Bottleneck Networks for a Person with Severe Hearing Loss
复制标题

DOI:
10.2197/ipsjtcva.7.64
复制
发表时间:
2015
期刊:
IPSJ Trans. Comput. Vis. Appl.
影响因子:
--
通讯作者:
Yuki Takashima;Yasuhiro Kakihara;Ryo Aihara;T. Takiguchi;Y. Ariki;Nobuyuki Mitani;K. Omori;Kaoru Nakazono
Yuki Takashima;Yasuhiro Kakihara;Ryo Aihara;T. Takiguchi;Y. Ariki;Nobuyuki Mitani;K. Omori;Kaoru Nakazono
中科院分区:
其他
文献类型:
--
作者:
Yuki Takashima;Yasuhiro Kakihara;Ryo Aihara;T. Takiguchi;Y. Ariki;Nobuyuki Mitani;K. Omori;Kaoru Nakazono

文献摘要

相似文献

在本文中,我们提出了一种视听语音识别系统,用于严重听力损失导致的发音障碍患者。对于患有这种发音障碍的人来说,他们的说话风格与没有听力损失的人的说话风格大不相同,因此对于没有听力损失的人来说,独立于说话者的模型几乎无法识别。本文研究了一种针对重度听力损失患者在噪声环境下的视听语音识别系统,将一种基于卷积瓶颈网络(CBN)的鲁棒特征提取方法应用于视听数据。我们通过噪声环境下的词识别实验验证了该方法的有效性,其中基于cbn的特征提取方法优于传统方法。
In this paper, we propose an audio-visual speech recognition system for a person with an articulation disorder resulting from severe hearing loss. In the case of a person with this type of articulation disorder, the speech style is quite different from with the result that of people without hearing loss that a speaker-independent model for unimpaired persons is hardly useful for recognizing it. We investigate in this paper an audio-visual speech recognition system for a person with severe hearing loss in noisy environments, where a robust feature extraction method using a convolutive bottleneck network (CBN) is applied to audio-visual data. We confirmed the effectiveness of this approach through word-recognition experiments in noisy environments, where the CBN-based feature extraction method outperformed the conventional methods.