Feature Frame Stacking in RNN-Based Tandem ASR Systems - Learned vs. Predefined Context

Feature Frame Stacking in RNN-Based Tandem ASR Systems - Learned vs. Predefined Context
复制标题

基于 RNN 的串联 ASR 系统中的特征帧堆叠 - 学习上下文与预定义上下文

DOI:
--
复制
发表时间:
2011
期刊:
Interspeech
影响因子:
--
通讯作者:
G. Rigoll
G. Rigoll
中科院分区:
--
文献类型:
--
作者:
M. Wöllmer;Björn Schuller;G. Rigoll

文献摘要

被引文献

相似文献

众所周知,音素识别得益于考虑上下文信息的技术,因此应用于串联自动语音识别(ASR)系统的神经网络通常采用某种形式的上下文建模。虽然基于多层感知器或递归神经网络(RNN)的方法能够通过同时处理连续特征向量的堆叠序列来对预定义量的上下文进行建模,但双向长短期记忆(BLSTM)网络被证明非常适合于将自学量的上下文用于音素预测。在本文中,我们评估了BLSTM建模和帧堆叠的组合,以确定在基于RNN的Tandem系统中利用上下文的最有效方法。应用CO-SINE语料库和我们最近引入的多流BLSTM-HMM解码器,我们为BLSTM网络冗余帧堆叠而RNN从预定义的特征级上下文中获利的直觉提供了经验证据。
As phoneme recognition is known to profit from techniques that consider contextual information, neural networks applied in Tandem automatic speech recognition (ASR) systems usually employ some form of context modeling. While approaches based on multi-layer perceptrons or recurrent neural networks (RNN) are able to model a predefined amount of context by simultaneously processing a stacked sequence of successive feature vectors, bidirectional Long Short-Term Memory (BLSTM) networks were shown to be well-suited for incorporating a self-learned amount of context for phoneme prediction. In this paper, we evaluate combinations of BLSTM modeling and frame stacking to determine the most efficient method for exploiting context in RNN-based Tandem systems. Applying the CO-SINE corpus and our recently introduced multi-stream BLSTM-HMM decoder, we provide empirical evidence for the intuition that BLSTM networks redundantize frame stacking while RNNs profit from predefined feature-level context.