Discriminative piecewise linear transformation based on deep learning for noise robust automatic speech recognition

Discriminative piecewise linear transformation based on deep learning for noise robust automatic speech recognition
复制标题

DOI:
10.1109/asru.2013.6707755
复制
发表时间:
2013-12
期刊:
2013 IEEE Workshop on Automatic Speech Recognition and Understanding
影响因子:
--
通讯作者:
Yosuke Kashiwagi;D. Saito;N. Minematsu;K. Hirose
Yosuke Kashiwagi;D. Saito;N. Minematsu;K. Hirose
中科院分区:
其他
文献类型:
--
作者:
Yosuke Kashiwagi;D. Saito;N. Minematsu;K. Hirose

文献摘要

相似文献

在本文中,我们提出使用深度神经网络来扩展基于分段线性变换的统计特征增强的传统方法。基于立体的分段线性环境补偿 (SPLICE) 是一种强大的特征增强统计方法,它将输入噪声特征的概率分布建模为高斯混合。然而,输入向量到划分区域的软分配有时做得不充分,并且向量会经历不充分的转换。特别是当转换必须是线性时,转换性能很容易下降。使用神经网络进行特征增强是另一种强大的方法,它可以直接对噪声和干净特征空间之间的非线性关系进行建模。然而,在这种情况下,它往往会出现过度拟合的问题。在本文中,我们尝试通过减少要估计的模型参数的数量来缓解这个问题。我们的神经网络经过训练,其输出层与干净特征空间中的状态相关联,而不是与噪声特征空间中的状态相关联。这种策略使得输出层的大小独立于给定噪声环境的类型。首先,我们将干净特征的分布描述为高斯混合模型,然后通过使用深度神经网络,有区别地估计输入噪声特征对应的干净空间中的状态。使用 Aurora 2 数据集的实验评估表明,与传统方法相比,我们提出的方法具有最佳性能。
In this paper, we propose the use of deep neural networks to expand conventional methods of statistical feature enhancement based on piecewise linear transformation. Stereo-based piecewise linear compensation for environments (SPLICE), which is a powerful statistical approach for feature enhancement, models the probabilistic distribution of input noisy features as a mixture of Gaussians. However, soft assignment of an input vector to divided regions is sometimes done inadequately and the vector comes to go through inadequate conversion. Especially when conversion has to be linear, the conversion performance will be easily degraded. Feature enhancement using neural networks is another powerful approach which can directly model a non-linear relationship between noisy and clean feature spaces. In this case, however, it tends to suffer from over-fitting problems. In this paper, we attempt to mitigate this problem by reducing the number of model parameters to estimate. Our neural network is trained whose output layer is associated with the states in the clean feature space, not in the noisy feature space. This strategy makes the size of the output layer independent of the kind of a given noisy environment. Firstly, we characterize the distribution of clean features as a Gaussian mixture model and then, by using deep neural networks, estimate discriminatively the state in the clean space that an input noisy feature corresponds to. Experimental evaluations using the Aurora 2 dataset demonstrate that our proposed method has the best performance compared to conventional methods.