Feature Frame Stacking in RNN-Based Tandem ASR Systems - Learned vs. Predefined Context
Feature Frame Stacking in RNN-Based Tandem ASR Systems - Learned vs. Predefined Context
复制标题
基于 RNN 的串联 ASR 系统中的特征帧堆叠 - 学习上下文与预定义上下文
DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
G. Rigoll
中科院分区:
文献类型:
--
作者:
M. Wöllmer;Björn Schuller;G. Rigoll
As phoneme recognition is known to profit from techniques that consider contextual information, neural networks applied in Tandem automatic speech recognition (ASR) systems usually employ some form of context modeling. While approaches based on multi-layer perceptrons or recurrent neural networks (RNN) are able to model a predefined amount of context by simultaneously processing a stacked sequence of successive feature vectors, bidirectional Long Short-Term Memory (BLSTM) networks were shown to be well-suited for incorporating a self-learned amount of context for phoneme prediction. In this paper, we evaluate combinations of BLSTM modeling and frame stacking to determine the most efficient method for exploiting context in RNN-based Tandem systems. Applying the CO-SINE corpus and our recently introduced multi-stream BLSTM-HMM decoder, we provide empirical evidence for the intuition that BLSTM networks redundantize frame stacking while RNNs profit from predefined feature-level context.