A new uncertainty decoding scheme for DNN-HMM hybrid systems with multichannel speech enhancement
A new uncertainty decoding scheme for DNN-HMM hybrid systems with multichannel speech enhancement
复制标题
DOI:
10.1109/icassp.2016.7472781
复制
发表时间:
2016-03
期刊:
影响因子:
--
通讯作者:
Christian Huemmer;A. Schwarz;R. Maas;Hendrik Barfuss;Ramón Fernández Astudillo;Walter Kellermann
中科院分区:
文献类型:
--
作者:
Christian Huemmer;A. Schwarz;R. Maas;Hendrik Barfuss;Ramón Fernández Astudillo;Walter Kellermann
Uncertainty decoding combines a probabilistic feature description with the acoustic model of a speech recognition system. For DNN-HMM hybrid systems, this can be realized by averaging the DNN outputs produced by a finite set of feature samples (drawn from an estimated probability distribution). In this article, we employ this sampling approach in combination with a multi-microphone speech enhancement system. We propose a new strategy for generating feature samples from multichannel signals, based on modeling the spatial coherence estimates between different microphone pairs as realizations of a latent random variable. From each coherence estimate, a spectral enhancement gain is computed and an enhanced feature vector is obtained, thus producing a finite set of feature samples, of which we average the respective DNN outputs. In the experimental part, this new uncertainty decoding strategy is shown to consistently improve the recognition accuracy of a DNN-HMM hybrid system for the 8-channel REVERB Challenge task.