Multi-Condition Training of Denoising Autoencoder by Augmenting Simulated Reverberant Speech Data
Multi-Condition Training of Denoising Autoencoder by Augmenting Simulated Reverberant Speech Data
复制标题
DOI:
10.1109/gcce.2018.8574776
复制
发表时间:
2018-10
期刊:
影响因子:
--
通讯作者:
R. Nahar;Takashi Kawai;A. Kai
中科院分区:
文献类型:
--
作者:
R. Nahar;Takashi Kawai;A. Kai
When speech recognition task is done by an automatic speech recognition system (ASR), it always has to process the noise and reverberation mixed with the speech in natural environment, which is a challenge. Recently, DNN-based ASR systems are showing better performance than other methods. Also, increasing training data for the DNN has been proven effective. But in real world, there can be many exceptions when it comes to collect data for training the DNN. In that case, if the characteristics of reverberant data for which the performance increases can be found, it will be a great help in the field of speech recognition. In this research, we analyze the effect of artificial reverberation condition on the data those are used for training DNN-based ASR front-end. There are possibilities to enhance the dataset by using suitable artificial data created by adding simulated reverberation to clean speech instead of using just real noisy recordings for training dataset. In case of simulated reverberation, the range of room impulse response (RIR) and direct-to-reverberant energy ratio (DRR) play important roles in training the system affecting the accuracy of recognition. Our analysis in this paper discusses the way of finding optimum RIR conditions to simulate the data without increasing training data amount.