Low Latency Acoustic Modeling Using Temporal Convolution and LSTMs

Low Latency Acoustic Modeling Using Temporal Convolution and LSTMs
复制标题

DOI:
10.1109/lsp.2017.2723507
复制
发表时间:
2018-03-01
影响因子:
3.9
通讯作者:
Khudanpur, Sanjeev
Khudanpur, Sanjeev
中科院分区:
工程技术2区
文献类型:
--
作者:
Peddinti, Vijayaditya;Wang, Yiming;Khudanpur, Sanjeev

文献摘要

被引文献

相似文献

Bidirectional long short-term memory (BLSTM) acoustic models provide a significant word error rate reduction compared to their unidirectional counterpart, as they model both the past and future temporal contexts. However, it is nontrivial to deploy bidirectional acoustic models for online speech recognition due to an increase in latency. In this letter, we propose the use of temporal convolution, in the form of time-delay neural network (TDNN) layers, along with unidirectional LSTM layers to limit the latency to 200 ms. This architecture has been shown to outperform the state-of-the-art low frame rate (LFR) BLSTM models. We further improve these LFR BLSTM acoustic models by operating them at higher frame rates at lower layers and show that the proposed model performs similar to these mixed frame rate BLSTMs. We present results on the Switchboard 300 h LVCSR task and the AMI LVCSR task, in the three microphone conditions.