Audiovisual Speech Activity Detection with Advanced Long Short-Term Memory

Audiovisual Speech Activity Detection with Advanced Long Short-Term Memory
复制标题

DOI:
10.21437/interspeech.2018-2490
复制
发表时间:
2018-09
期刊:
--
影响因子:
--
通讯作者:
Fei Tao;C. Busso
Fei Tao;C. Busso
中科院分区:
其他
文献类型:
--
作者:
Fei Tao;C. Busso

文献摘要

被引文献

相似文献

语音活动检测 (SAD) 是基于语音的系统的关键预处理步骤。传统的纯音频 SAD (A-SAD) 系统在实际应用中会受到噪声的影响。解决此问题的另一种方法是包含视觉信息,创建视听语音活动检测 (AV-SAD) 解决方案。在我们之前的工作中,我们提出使用双模循环神经网络(BRNN)构建 AV-SAD 系统。该框架能够捕获音频和视觉输入中与任务相关的特征,并对模态内部和跨模态的时间信息进行建模。该方法依赖于长短期记忆(LSTM)。尽管 LSTM 可以对单元的较长时间依赖性进行建模,但单元的有效记忆仅限于几个帧,因为循环连接仅考虑前一帧。对于 SAD 系统,重要的是对较长的时间依赖性进行建模,以捕获声学和口面部特征中传达的语音的半周期性质。本研究提出使用先进的 LSTM (A-LSTM) 实现基于 BRNN 的 AV-SAD 系统,该系统通过在过去包含与帧的多个连接来克服这一限制。结果表明,所提出的框架可以显着优于使用原始 LSTM 层训练的 BRNN 系统。
Speech activity detection (SAD) is a key pre-processing step for a speech-based system. The performance of conventional audio-only SAD (A-SAD) systems is impaired by acoustic noise when they are used in practical applications. An alternative approach to address this problem is to include visual information, creating audiovisual speech activity detection (AV-SAD) solutions. In our previous work, we proposed to build an AV-SAD system using bimodal recurrent neural network (BRNN). This framework was able to capture the task-related characteristics in the audio and visual inputs, and model the temporal information within and across modalities. The approach relied on long short-term memory (LSTM). Although LSTM can model longer temporal dependencies with the cells, the effective memory of the units is limited to a few frames, since the recurrent connection only considers the previous frame. For SAD systems, it is important to model longer temporal dependencies to capture the semi-periodic nature of speech conveyed in acoustic and orofacial features. This study proposes to implement a BRNN-based AV-SAD system with advanced LSTMs (A-LSTMs), which overcomes this limitation by including multiple connections to frames in the past. The results show that the proposed framework can significantly outperform the BRNN system trained with the original LSTM layers.