Effect of Data Augmentation on DNN-Based VAD for Automatic Speech Recognition in Noisy Environment

Effect of Data Augmentation on DNN-Based VAD for Automatic Speech Recognition in Noisy Environment
复制标题

DOI:
10.1109/gcce50665.2020.9292034
复制
发表时间:
2020-10
期刊:
2020 IEEE 9th Global Conference on Consumer Electronics (GCCE)
影响因子:
--
通讯作者:
R. Nahar;A. Kai
R. Nahar;A. Kai
中科院分区:
其他
文献类型:
--
作者:
R. Nahar;A. Kai

文献摘要

相似文献

近年来,基于DNN的系统非常流行,用于执行自动语音识别(ASR)任务。它们表现出了比其他方法更好的性能。为了有效地执行ASR任务,输入数据的分割也必须准确。有不同类型的语音活动检测(VAD)方法,包括基于功率的统计方法。但要与ASR系统结合,最好使用基于DNN的方法来执行VAD任务。在本文中,我们研究了基于DNN的VAD方法相对于基于功率的统计方法的优越性。不仅如此,还研究了在不同条件下数据增强对VAD性能的影响。使用专门为评估VAD任务而开发的CENSREC-1-C数据来评估VAD的性能。使用多条件数据训练的VAD表现出比基准和其他流行的基于功率的统计VAD工具更好的性能。我们还评估了CSJ语料库及其电话变体的val1测试数据集的性能。
In recent years, DNN-based systems are extremely popular for performing automatic speech recognition (ASR) task. They have shown better performance than other methods. To perform ASR task efficiently, segmentation of the input data also has to be accurate. There are different kinds of methods of voice activity detection (VAD) including power based statistical methods. But to incorporate with the ASR system, it is better to perform the VAD task with the help of DNN-based method. In this paper, we investigate the superiority of DNN-based VAD over power based statistical methods. Not only that but also the effect of data augmentation on the performance of VAD in various conditions is also investigated. Performance of the VAD is evaluated using CENSREC-1-C data especially developed for evaluation of VAD task. The VAD trained with multi-condition data shows better performance than the baseline and other popular power based statistical VAD tools. We have also evaluated the performance for eval1 test dataset of CSJ corpus as well as its telephone variants.