Self-attention Aligner: A Latency-control End-to-end Model for ASR Using Self-attention Network and Chunk-hopping

Self-attention Aligner: A Latency-control End-to-end Model for ASR Using Self-attention Network and Chunk-hopping
复制标题

DOI:
10.1109/icassp.2019.8682954
复制
发表时间:
2019-02
期刊:
ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Linhao Dong;Feng Wang;Bo Xu
Linhao Dong;Feng Wang;Bo Xu
中科院分区:
其他
文献类型:
--
作者:
Linhao Dong;Feng Wang;Bo Xu

文献摘要

被引文献

相似文献

自注意力网络是一种基于注意力的前馈神经网络,最近显示出在各种NLP任务中取代递归神经网络(RNN)的潜力。然而,目前尚不清楚自注意力网络是否可以作为自动语音识别(ASR)中RNN的良好替代方案,它处理较长的语音序列,并且可能具有在线识别要求。在本文中,我们提出了一个无RNN的端到端模型:自注意力对齐器(SAA),它将自注意力网络应用于简化的递归神经对齐器(RNA)框架。我们还提出了一个组块跳跃机制,这使得SAA模型的编码分割帧块一个接一个,以支持在线识别。在两个普通话ASR数据集上的实验表明,用自注意网络代替RNN,相对字符错误率(CER)降低了8.4%-10.2%。此外,组块跳跃机制允许SAA具有仅2.5%的相对CER降级和320ms的延迟。在与自注意网络语言模型联合训练后,我们的SAA模型在多个数据集上获得了进一步的错误率降低。特别是,它在普通话ASR基准(HKUST)上实现了24.12%的CER,超过了最好的端到端模型超过2%的绝对CER。
Self-attention network, an attention-based feedforward neural network, has recently shown the potential to replace recurrent neural networks (RNNs) in a variety of NLP tasks. However, it is not clear if the self-attention network could be a good alternative of RNNs in automatic speech recognition (ASR), which processes the longer speech sequences and may have online recognition requirements. In this paper, we present a RNN-free end-to-end model: self-attention aligner (SAA), which applies the self-attention networks to a simplified recurrent neural aligner (RNA) framework. We also propose a chunk-hopping mechanism, which enables the SAA model to encode on segmented frame chunks one after another to support online recognition. Experiments on two Mandarin ASR datasets show the replacement of RNNs by the self-attention networks yields a 8.4%-10.2% relative character error rate (CER) reduction. In addition, the chunk-hopping mechanism allows the SAA to have only a 2.5% relative CER degradation with a 320ms latency. After jointly training with a self-attention network language model, our SAA model obtains further error rate reduction on multiple datasets. Especially, it achieves 24.12% CER on the Mandarin ASR benchmark (HKUST), exceeding the best end-to-end model by over 2% absolute CER.