An Investigation of Enhancing CTC Model for Triggered Attention-based Streaming ASR

An Investigation of Enhancing CTC Model for Triggered Attention-based Streaming ASR
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
影响因子:
--
通讯作者:
Huaibo Zhao;Yosuke Higuchi;Tetsuji Ogawa;Tetsunori Kobayashi
Huaibo Zhao;Yosuke Higuchi;Tetsuji Ogawa;Tetsunori Kobayashi
中科院分区:
其他
文献类型:
--
作者:
Huaibo Zhao;Yosuke Higuchi;Tetsuji Ogawa;Tetsunori Kobayashi

文献摘要

相似文献

本文尝试将联合收割机Mask-CTC和触发注意机制相结合,构建一个具有高性能和低延迟的流媒体端到端自动语音识别(ASR)系统。执行由CTC尖峰触发的自回归解码的触发注意机制已被证明在流式ASR中是有效的。然而,为了保持基于CTC输出的对准估计的高精度(这是其性能的关键),不可避免的是,应当利用一些未来信息输入(即,具有更高的延迟)。应当注意,在流式传输ASR中,期望能够实现高识别准确度,同时保持低延迟。因此,本研究旨在通过引入Mask-CTC来实现具有低延迟的高度准确的流式ASR,Mask-CTC能够学习预测未来信息的特征表示(即,可以考虑长期上下文)到编码器预训练。使用华尔街日报数据进行的实验比较表明,该方法实现了更高的准确性与更低的延迟比传统的触发注意力为基础的流ASR系统。
In the present paper, an attempt is made to combine Mask-CTC and the triggered attention mechanism to construct a streaming end-to-end automatic speech recognition (ASR) system that provides high performance with low latency. The triggered attention mechanism, which performs autoregressive decoding triggered by the CTC spike, has shown to be effective in streaming ASR. However, in order to maintain high accuracy of alignment estimation based on CTC outputs, which is the key to its performance, it is inevitable that decoding should be performed with some future information input (i.e., with higher latency). It should be noted that in streaming ASR, it is desirable to be able to achieve high recognition accuracy while keeping the latency low. Therefore, the present study aims to achieve highly accurate streaming ASR with low latency by introducing Mask-CTC, which is capable of learning feature representations that anticipate future information (i.e., that can consider long-term contexts), to the encoder pre-training. Experimental comparisons conducted using WSJ data demonstrate that the proposed method achieves higher accuracy with lower latency than the conventional triggered attention-based streaming ASR system.