A new uncertainty decoding scheme for DNN-HMM hybrid systems with multichannel speech enhancement

A new uncertainty decoding scheme for DNN-HMM hybrid systems with multichannel speech enhancement
复制标题

DOI:
10.1109/icassp.2016.7472781
复制
发表时间:
2016-03
期刊:
2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Christian Huemmer;A. Schwarz;R. Maas;Hendrik Barfuss;Ramón Fernández Astudillo;Walter Kellermann
Christian Huemmer;A. Schwarz;R. Maas;Hendrik Barfuss;Ramón Fernández Astudillo;Walter Kellermann
中科院分区:
其他
文献类型:
--
作者:
Christian Huemmer;A. Schwarz;R. Maas;Hendrik Barfuss;Ramón Fernández Astudillo;Walter Kellermann

文献摘要

被引文献

相似文献

不确定性解码将概率特征描述与语音识别系统的声学模型相结合。对于DNN-HMM混合系统,这可以通过对有限的特征样本集(从估计的概率分布中提取)产生的DNN输出进行平均来实现。在本文中,我们将这种采样方法与多麦克风语音增强系统相结合。我们提出了一种新的从多通道信号中生成特征样本的策略,该策略基于将不同麦克风对之间的空间相干性估计建模为潜在随机变量的实现。根据每个相干估计,计算频谱增强增益并获得增强的特征向量,从而产生有限的特征样本集,我们对其各自的DNN输出进行平均。在实验部分,这种新的不确定性解码策略持续提高了DNN-HMM混合系统对8声道混响挑战任务的识别精度。
Uncertainty decoding combines a probabilistic feature description with the acoustic model of a speech recognition system. For DNN-HMM hybrid systems, this can be realized by averaging the DNN outputs produced by a finite set of feature samples (drawn from an estimated probability distribution). In this article, we employ this sampling approach in combination with a multi-microphone speech enhancement system. We propose a new strategy for generating feature samples from multichannel signals, based on modeling the spatial coherence estimates between different microphone pairs as realizations of a latent random variable. From each coherence estimate, a spectral enhancement gain is computed and an enhanced feature vector is obtained, thus producing a finite set of feature samples, of which we average the respective DNN outputs. In the experimental part, this new uncertainty decoding strategy is shown to consistently improve the recognition accuracy of a DNN-HMM hybrid system for the 8-channel REVERB Challenge task.