Probabilistic speech feature extraction with context-sensitive Bottleneck neural networks

Probabilistic speech feature extraction with context-sensitive Bottleneck neural networks
复制标题

使用上下文敏感瓶颈神经网络进行概率语音特征提取

DOI:
10.1016/j.neucom.2012.06.064
复制
发表时间:
2014
期刊:
影响因子:
6
通讯作者:
B. Schuller
B. Schuller
中科院分区:
计算机科学2区
文献类型:
--
作者:
M. Wöllmer;B. Schuller

文献摘要

参考文献

被引文献

相似文献

本文提出了一种新的上下文敏感的特征提取方法,用于自发语音识别。由于双向长短期记忆(BLSTM)网络是已知的,使提高音素识别的准确性,通过将远程上下文信息到语音解码,我们集成的BLSTM原则到串联前端的概率特征提取。与以前提出的方法,利用BLSTM建模生成一个离散的音素预测功能,我们的特征提取器合并连续的高级别的概率BLSTM功能与低级别的功能。通过结合BLSTM建模和瓶颈(BN)特征生成,我们提出了一种新的前端,使我们能够产生任意大小的上下文敏感的概率特征向量,独立于网络训练目标。具有挑战性的自发,会话语音识别任务的评估表明,这个概念优于最近发布的架构特征级上下文建模。
We introduce a novel context-sensitive feature extraction approach for spontaneous speech recognition. As bidirectional Long Short-Term Memory (BLSTM) networks are known to enable improved phoneme recognition accuracies by incorporating long-range contextual information into speech decoding, we integrate the BLSTM principle into a Tandem front-end for probabilistic feature extraction. Unlike the previously proposed approaches which exploit BLSTM modeling by generating a discrete phoneme prediction feature, our feature extractor merges continuous high-level probabilistic BLSTM features with low-level features. By combining BLSTM modeling and Bottleneck (BN) feature generation, we propose a novel front-end that allows us to produce context-sensitive probabilistic feature vectors of arbitrary size, independent of the network training targets. Evaluations on challenging spontaneous, conversational speech recognition tasks show that this concept prevails over recently published architectures for feature-level context modeling.
多流 HMM/ANN 混合方法实现噪声鲁棒 ASR 的最新进展
DOI: --
发表时间: 2005
影响因子: 4.3
作者:
Astrid Hagen;A. Morris
通讯作者: A. Morris
基于 RNN 的串联 ASR 系统中的特征帧堆叠 - 学习上下文与预定义上下文
DOI: --
发表时间: 2011
期刊: Interspeech
影响因子: --
作者:
M. Wöllmer;Björn Schuller;G. Rigoll
通讯作者: G. Rigoll
DOI: --
发表时间: 2004
期刊: Machine Learning for Multimodal Interaction
影响因子: --
作者:
Q. Zhu;Barry Y. Chen;N. Morgan;A. Stolcke
通讯作者: A. Stolcke
DOI: --
发表时间: 2010
期刊: Interspeech
影响因子: --
作者:
M. Wöllmer;F. Eyben;Björn Schuller;G. Rigoll
通讯作者: G. Rigoll
噪声语音识别:鲁棒模型架构和特征增强的比较调查
DOI: --
发表时间: 2009
期刊: EURASIP Journal on Audio, Speech, and Music Processing
影响因子: --
作者:
Björn Schuller;M. Wöllmer;T. Moosmayr;G. Rigoll
通讯作者: G. Rigoll