Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR

Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR
复制标题

DOI:
10.1109/icassp40776.2020.9054098
复制
发表时间:
2020-04
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
H. Inaguma;Yashesh Gaur;Liang Lu;Jinyu Li;Y. Gong
H. Inaguma;Yashesh Gaur;Liang Lu;Jinyu Li;Y. Gong
中科院分区:
其他
文献类型:
--
作者:
H. Inaguma;Yashesh Gaur;Liang Lu;Jinyu Li;Y. Gong

文献摘要

相似文献

最近,一些新的流注意力为基础的序列到序列(S2S)模型已被提出来执行线性时间解码复杂度的在线语音识别。然而,在这些模型中,与实际声学边界相比,生成令牌的决策被延迟,因为它们的单向编码器缺乏未来信息。这导致推理过程中不可避免的延迟。为了缓解这个问题并减少延迟,我们在训练过程中提出了几种策略,利用从混合模型中提取的外部硬对齐。我们调查,利用在编码器和解码器的对齐。在编码器侧,研究了(1)多任务学习和(2)使用逐帧分类任务的预训练。在解码器侧,我们(3)在对齐边缘化期间移除超过可接受延迟的不适当对齐路径,以及(4)直接最小化可区分的预期延迟损失。在Cortana语音搜索任务上的实验表明,我们提出的方法可以显着减少延迟,甚至在某些情况下在解码器端提高识别精度。我们还提出了一些分析,以了解流S2S模型的行为。
Recently, a few novel streaming attention-based sequence-to-sequence (S2S) models have been proposed to perform online speech recognition with linear-time decoding complexity. However, in these models, the decisions to generate tokens are delayed compared to the actual acoustic boundaries since their unidirectional encoders lack future information. This leads to an inevitable latency during inference. To alleviate this issue and reduce latency, we propose several strategies during training by leveraging external hard alignments extracted from the hybrid model. We investigate to utilize the alignments in both the encoder and the decoder. On the encoder side, (1) multi-task learning and (2) pre-training with the framewise classification task are studied. On the decoder side, we (3) remove inappropriate alignment paths beyond an acceptable latency during the alignment marginalization, and (4) directly min-imize the differentiable expected latency loss. Experiments on the Cortana voice search task demonstrate that our proposed methods can significantly reduce the latency, and even improve the recognition accuracy in certain cases on the decoder side. We also present some analysis to understand the behaviors of streaming S2S models.