Recognition of spontaneous conversational speech using long short-term memory phoneme predictions
Recognition of spontaneous conversational speech using long short-term memory phoneme predictions
复制标题
使用长短期记忆音素预测识别自发会话语音
DOI:
--
复制
发表时间:
2010
期刊:
影响因子:
--
通讯作者:
G. Rigoll
中科院分区:
文献类型:
--
作者:
M. Wöllmer;F. Eyben;Björn Schuller;G. Rigoll
We present a novel continuous speech recognition framework designed to unite the principles of triphone and Long ShortTerm Memory (LSTM) modeling. The LSTM principle allows a recurrent neural network to store and to retrieve information over long time periods, which was shown to be well-suited for the modeling of co-articulation effects in human speech. Our system uses a bidirectional LSTM network to generate a phoneme prediction feature that is observed by a triphone-based large-vocabulary continuous speech recognition (LVCSR) decoder, together with conventional MFCC features. We evaluate both, phoneme prediction error rates of various network architectures and the word recognition performance of our Tandem approach using the COSINE database - a large corpus of conversational and noisy speech, and show that incorporating LSTM phoneme predictions in to an LVCSR system leads to significantly higher word accuracies. Index Terms: Long Short-Term Memory, Large-Vocabulary Continuous Speech Recognition, Context Modeling, Recurrent Neural Networks