Reverberant speech recognition based on denoising autoencoder

Reverberant speech recognition based on denoising autoencoder
复制标题

DOI:
10.21437/interspeech.2013-267
复制
发表时间:
2013
期刊:
--
影响因子:
--
通讯作者:
Takaaki Ishii;Hiroki Komiyama;T. Shinozaki;Y. Horiuchi;S. Kuroiwa
Takaaki Ishii;Hiroki Komiyama;T. Shinozaki;Y. Horiuchi;S. Kuroiwa
中科院分区:
其他
文献类型:
--
作者:
Takaaki Ishii;Hiroki Komiyama;T. Shinozaki;Y. Horiuchi;S. Kuroiwa

文献摘要

被引文献

相似文献

将去噪自编码器应用于混响语音识别中,作为抗噪前端,从含噪输入中重建出清晰的语音频谱。为了捕获语音声音的上下文效果,多个短窗口频谱帧的窗口被级联以形成单个输入向量。此外,短期和长期频谱的组合进行了研究,以适当地处理混响的长脉冲响应,同时保持必要的时间分辨率的语音识别。使用CENSREC-4数据集进行实验,该数据集被设计为远距离说话语音识别的评估框架。实验结果表明,所提出的去噪自动编码器的前端使用的短波频谱给出了更好的效果比传统的方法。通过结合长期光谱,获得了进一步的改进。所提出的方法使用短期和长期光谱的识别准确率为97.0%的开放条件测试集的数据集,而它是87.8%时,使用多条件训练为基础的基线。作为补充实验,大词汇量的语音识别和所提出的方法已被证实的有效性。索引术语:去噪自动编码器,混响语音识别,受限玻尔兹曼机,远距离语音识别,CENSREC-4
Denoising autoencoder is applied to reverberant speech recognition as a noise robust front-end to reconstruct clean speech spectrum from noisy input. In order to capture context effects of speech sounds, a window of multiple short-windowed spectral frames are concatenated to form a single input vector. Additionally, a combination of short and long-term spectra is investigated to properly handle long impulse response of reverberation while keeping necessary time resolution for speech recognition. Experiments are performed using the CENSREC-4dataset that is designed as an evaluation framework for distant-talking speech recognition. Experimental results show that the proposed denoising autoencoder based front-end using the shortwindowed spectra gives better results than conventional methods. By combining the long-term spectra, further improvement is obtained. The recognition accuracy by the proposed method using the short and long-term spectra is 97.0% for the open condition test set of the dataset, whereas it is 87.8% when a multicondition training based baseline is used. As a supplemental experiment, large vocabulary speech recognition is also performed and the effectiveness of the proposed method has been confirmed. Index Terms: Denoising autoencoder, reverberant speech recognition, restricted Boltzmann machine, distant-talking speech recognition, CENSREC-4