Increasing the robustness of CNN acoustic models using autoregressive moving average spectrogram features and channel dropout

Increasing the robustness of CNN acoustic models using autoregressive moving average spectrogram features and channel dropout
复制标题

DOI:
10.1016/j.patrec.2017.09.023
复制
发表时间:
2017-12
期刊:
Pattern Recognit. Lett.
影响因子:
--
通讯作者:
György Kovács;L. Tóth;Dirk Van Compernolle;Sriram Ganapathy
György Kovács;L. Tóth;Dirk Van Compernolle;Sriram Ganapathy
中科院分区:
其他
文献类型:
--
作者:
György Kovács;L. Tóth;Dirk Van Compernolle;Sriram Ganapathy

文献摘要

被引文献

相似文献

开发对不匹配和噪声信道条件具有鲁棒性的自动语音识别系统是一个具有挑战性的问题,特别是在训练和测试条件不同的情况下。在这里,我们试图通过结合两种方法来增加卷积神经网络(CNN)声学模型在这种情况下的鲁棒性。首先,我们利用输入时频表示的特殊结构,提出了一种改进版本的输入dropout。所提出的信道丢弃方法抛弃了整个频谱信道,而不是仅仅丢弃频谱图的随机“像素”。我们期望这种放弃策略将迫使网络减少对整个频谱的依赖,并使其对信道失配和窄带噪声具有更强的鲁棒性。其次,我们用自回归移动平均(ARMA)谱图取代了标准的梅尔谱图输入表示,该自回归移动平均(ARMA)谱图最近被证明在不匹配的训练测试条件下优于前者。在我们对Aurora-4数据库的实验中,当使用干净的训练场景时,与基线CNN相比,所提出的通道丢弃方法在ARMA特征上的相对词错误率降低了16%(绝对提高了3%),在FBANK特征上的相对词错误率降低了20%(绝对提高了7%)。
Developing automatic speech recognition systems that are robust to mismatched and noisy channel conditions is a challenging problem, especially when the training and the test conditions are different. Here, we seek to increase the robustness of convolutional neural network (CNN) acoustic models under such circumstances by combining two methods. Firstly, we propose an improved version of input dropout, which exploits the special structure of the input time-frequency representation. Instead of just dropping out random ‘pixels’ of the spectrogram, the proposed channel dropout approach discards whole spectral channels. We expect that this dropout strategy will force the network to rely less on the whole spectrum, and make it more robust to channel mismatches and narrow-band noise. Secondly, we replaced the standard mel-spectrogram input representation with the autoregressive moving average (ARMA) spectrogram, which was recently shown to outperform the former under mismatched train-test conditions. In our experiments on the Aurora-4 database, the proposed channel dropout method attained relative word error rate reductions of 16% with ARMA features (an absolute improvement of 3%), and 20% with FBANK features (an absolute improvement of 7%) over the baseline CNN, when using the clean training scenario.