Recognition of spontaneous conversational speech using long short-term memory phoneme predictions

Recognition of spontaneous conversational speech using long short-term memory phoneme predictions
复制标题

使用长短期记忆音素预测识别自发会话语音

DOI:
--
复制
发表时间:
2010
期刊:
Interspeech
影响因子:
--
通讯作者:
G. Rigoll
G. Rigoll
中科院分区:
--
文献类型:
--
作者:
M. Wöllmer;F. Eyben;Björn Schuller;G. Rigoll

文献摘要

被引文献

相似文献

我们提出了一种新的连续语音识别框架,旨在将三音子和长短期记忆(LSTM)建模的原理结合起来。LSTM原理允许递归神经网络在很长一段时间内存储和检索信息,这被证明非常适合人类语音中的协同发音效应建模。我们的系统使用双向LSTM网络来生成音素预测特征,该特征由基于三音素的大词汇量连续语音识别(LVCSR)解码器与传统的MFCC特征一起观察。我们使用COSINE数据库(一个大型的会话和嘈杂语音语料库)评估了各种网络架构的音素预测错误率和Tandem方法的单词识别性能,并表明将LSTM音素预测纳入LVCSR系统会导致显着更高的单词准确率。索引术语:长短期记忆,大词汇量连续语音识别,上下文建模,递归神经网络
We present a novel continuous speech recognition framework designed to unite the principles of triphone and Long ShortTerm Memory (LSTM) modeling. The LSTM principle allows a recurrent neural network to store and to retrieve information over long time periods, which was shown to be well-suited for the modeling of co-articulation effects in human speech. Our system uses a bidirectional LSTM network to generate a phoneme prediction feature that is observed by a triphone-based large-vocabulary continuous speech recognition (LVCSR) decoder, together with conventional MFCC features. We evaluate both, phoneme prediction error rates of various network architectures and the word recognition performance of our Tandem approach using the COSINE database - a large corpus of conversational and noisy speech, and show that incorporating LSTM phoneme predictions in to an LVCSR system leads to significantly higher word accuracies. Index Terms: Long Short-Term Memory, Large-Vocabulary Continuous Speech Recognition, Context Modeling, Recurrent Neural Networks