Speaker-Independent Silent Speech Recognition from Flesh-Point Articulatory Movements Using an LSTM Neural Network.

Speaker-Independent Silent Speech Recognition from Flesh-Point Articulatory Movements Using an LSTM Neural Network.
复制标题

DOI:
10.1109/taslp.2017.2758999
复制
发表时间:
2017-12
期刊:
IEEE/ACM transactions on audio, speech, and language processing
影响因子:
--
通讯作者:
Wang J
Wang J
中科院分区:
其他
文献类型:
--
作者:
Kim M;Cao B;Mau T;Wang J

文献摘要

被引文献

相似文献

无声语音识别(SSR)将非音频信息(如发音动作)转换为文本。SSR有可能使喉切除术患者通过自然的口语表达进行交流。目前的SSR系统在很大程度上依赖于说话人相关的识别模型。不同说话者之间发音模式的高度可变性已经成为开发有效的独立于说话者的SSR方法的障碍。然而,独立于说话人的SSR方法对于减少每个说话人所需的训练数据量至关重要。本文采用发音归一化方法,从舌和唇肉点的运动中研究与说话人无关的SSR,以减少说话人之间的差异。为了最大限度地减少跨说话人的发音生理差异,我们提出了基于Procrustes匹配的发音归一化方法,通过去除位置、旋转和缩放差异。为了进一步规范化衔接数据,我们应用特征空间最大似然线性回归和i向量。本文采用双向长短期记忆递归神经网络(BLSTM)作为发音模型,对具有长时间发音历史的发音运动进行有效建模。使用电磁关节记录仪(EMA)收集了12名健康和2名喉切除的英语使用者的无声语音数据集。实验结果表明,我们的独立于说话人的SSR方法对健康和喉切除术的说话人都是有效的。此外,BLSTM优于标准深度神经网络。三种归一化方法结合使用时,BLSTM的性能最好。
Silent speech recognition (SSR) converts non-audio information such as articulatory movements into text. SSR has the potential to enable persons with laryngectomy to communicate through natural spoken expression. Current SSR systems have largely relied on speaker-dependent recognition models. The high degree of variability in articulatory patterns across different speakers has been a barrier for developing effective speaker-independent SSR approaches. Speaker-independent SSR approaches, however, are critical for reducing the amount of training data required from each speaker. In this paper, we investigate speaker-independent SSR from the movements of flesh points on tongue and lip with articulatory normalization methods that reduce the inter-speaker variation. To minimize the across-speaker physiological differences of the articulators, we propose Procrustes matching-based articulatory normalization by removing locational, rotational, and scaling differences. To further normalize the articulatory data, we apply feature-space maximum likelihood linear regression and i-vector. In this paper, we adopt a bidirectional long short term memory recurrent neural network (BLSTM) as an articulatory model to effectively model the articulatory movements with long-range articulatory history. A silent speech data set with flesh points was collected using an electromagnetic articulograph (EMA) from twelve healthy and two laryngectomized English speakers. Experimental results showed the effectiveness of our speaker-independent SSR approaches on healthy as well as laryngectomy speakers. In addition, BLSTM outperformed standard deep neural network. The best performance was obtained by BLSTM with all the three normalization approaches combined.