Multi-Condition Training of Denoising Autoencoder by Augmenting Simulated Reverberant Speech Data

Multi-Condition Training of Denoising Autoencoder by Augmenting Simulated Reverberant Speech Data
复制标题

DOI:
10.1109/gcce.2018.8574776
复制
发表时间:
2018-10
期刊:
2018 IEEE 7th Global Conference on Consumer Electronics (GCCE)
影响因子:
--
通讯作者:
R. Nahar;Takashi Kawai;A. Kai
R. Nahar;Takashi Kawai;A. Kai
中科院分区:
其他
文献类型:
--
作者:
R. Nahar;Takashi Kawai;A. Kai

文献摘要

相似文献

自动语音识别系统(ASR)在进行语音识别任务时,往往要处理自然环境中语音中夹杂的噪声和混响,这是一个挑战。最近,基于DNN的ASR系统表现出比其他方法更好的性能。此外,增加DNN的训练数据已被证明是有效的。但在真实的世界中,当涉及到收集数据用于训练DNN时,可能会有许多例外。在这种情况下,如果能够找到性能提高的混响数据的特性,这将在语音识别领域中有很大的帮助。在这项研究中,我们分析了人工混响条件对用于训练基于DNN的ASR前端的数据的影响。有可能通过使用合适的人工数据来增强数据集,而不是仅使用真实的噪声记录来训练数据集,所述人工数据是通过将模拟混响添加到干净的语音而创建的。在模拟混响的情况下,房间脉冲响应(RIR)和直达混响能量比(DRR)的范围在训练系统中起着重要的作用,影响识别的准确性。本文的分析讨论了如何在不增加训练数据量的情况下找到最佳的RIR条件来模拟数据。
When speech recognition task is done by an automatic speech recognition system (ASR), it always has to process the noise and reverberation mixed with the speech in natural environment, which is a challenge. Recently, DNN-based ASR systems are showing better performance than other methods. Also, increasing training data for the DNN has been proven effective. But in real world, there can be many exceptions when it comes to collect data for training the DNN. In that case, if the characteristics of reverberant data for which the performance increases can be found, it will be a great help in the field of speech recognition. In this research, we analyze the effect of artificial reverberation condition on the data those are used for training DNN-based ASR front-end. There are possibilities to enhance the dataset by using suitable artificial data created by adding simulated reverberation to clean speech instead of using just real noisy recordings for training dataset. In case of simulated reverberation, the range of room impulse response (RIR) and direct-to-reverberant energy ratio (DRR) play important roles in training the system affecting the accuracy of recognition. Our analysis in this paper discusses the way of finding optimum RIR conditions to simulate the data without increasing training data amount.