Spatial diffuseness features for DNN-based speech recognition in noisy and reverberant environments

Spatial diffuseness features for DNN-based speech recognition in noisy and reverberant environments
复制标题

DOI:
10.1109/icassp.2015.7178798
复制
发表时间:
2014-10
期刊:
2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
A. Schwarz;Christian Huemmer;R. Maas;Walter Kellermann
A. Schwarz;Christian Huemmer;R. Maas;Walter Kellermann
中科院分区:
其他
文献类型:
--
作者:
A. Schwarz;Christian Huemmer;R. Maas;Walter Kellermann

文献摘要

被引文献

相似文献

为了在混响和噪声环境中提高识别精度,提出了一种基于深度神经网络(DNN)的自动语音识别的空间扩散特征。该特征是从多个麦克风信号实时计算的,不需要知道或估计到达方向,并且表示每个时间和频率区段中扩散噪声的相对量。结果表明,与从噪声信号中提取的对数谱特征相比,使用扩散度特征作为基于DNN的声学模型的附加输入可以降低混响挑战语料库的错误率,并且通过谱减法增强了特征。
We propose a spatial diffuseness feature for deep neural network (DNN)-based automatic speech recognition to improve recognition accuracy in reverberant and noisy environments. The feature is computed in real-time from multiple microphone signals without requiring knowledge or estimation of the direction of arrival, and represents the relative amount of diffuse noise in each time and frequency bin. It is shown that using the diffuseness feature as an additional input to a DNN-based acoustic model leads to a reduced word error rate for the REVERB challenge corpus, both compared to logmelspec features extracted from noisy signals, and features enhanced by spectral subtraction.