WaveNet With Cross-Attention for Audiovisual Speech Recognition

WaveNet With Cross-Attention for Audiovisual Speech Recognition
复制标题

DOI:
10.1109/access.2020.3024218
复制
发表时间:
2020
期刊:
影响因子:
3.9
通讯作者:
Hui Wang;Fei Gao;Yue Zhao;Licheng Wu
Hui Wang;Fei Gao;Yue Zhao;Licheng Wu
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hui Wang;Fei Gao;Yue Zhao;Licheng Wu

文献摘要

相似文献

本文提出了一种用于视听自动语音识别(AV-ASR)的交叉注意WaveNet,以解决两个数据流之间的多模态特征融合和帧对齐问题。WaveNet通常用于语音生成和语音识别,但本文将其扩展到视听语音识别,并在WaveNet的不同地方引入交叉注意机制进行特征融合。所提出的交叉注意机制试图探索视觉特征框架与听觉特征框架的相关框架。与传统的AV-ASR特征拼接方法相比,英文单词误差约为39.1%,英文单词误差约为21.6%。
In this paper, the WaveNet with cross-attention is proposed for Audio-Visual Automatic Speech Recognition (AV-ASR) to address multimodal feature fusion and frame alignment problems between two data streams. WaveNet is usually used for speech generation and speech recognition, however, in this paper, we extent it to audiovisual speech recognition, and the cross-attention mechanism is introduced into different places of WaveNet for feature fusion. The proposed cross-attention mechanism tries to explore the correlated frames of visual feature to the acoustic feature frame. The experimental results show that the WaveNet with cross-attention can reduce the Tibetan single syllable error about 4.5% and English word error about 39.8% relative to the audio-only speech recognition, and reduce Tibetan single syllable error about 35.1% and English word error about 21.6% relative to the conventional feature concatenation method for AV-ASR.