rVAD: An unsupervised segment-based robust voice activity detection method

rVAD: An unsupervised segment-based robust voice activity detection method
复制标题

DOI:
10.1016/j.csl.2019.06.005
复制
发表时间:
2020-01-01
影响因子:
4.3
通讯作者:
Dehak, Najim
Dehak, Najim
中科院分区:
计算机科学3区
文献类型:
--
作者:
Tan, Zheng-Hua;Sarkar, Achintya Kr;Dehak, Najim

文献摘要

被引文献

相似文献

本文提出了一种无监督的基于段的鲁棒语音活动检测(rVAD)方法。该方法包括两个通道的去噪,然后是语音活动检测(VAD)阶段。在第一遍中,通过使用后验信噪比(SNR)加权能量差来检测语音信号中的高能量段,并且如果在段内没有检测到音高,则该段被认为是高能量噪声段并且被设置为零。在第二次通过中,语音信号去噪的语音增强方法,其中几种方法进行了探索。接下来,具有音高的相邻帧被分组在一起以形成音高段,并且基于语音统计,音高段从两端进一步扩展,以便包括有声和无声声音以及可能的非语音部分。最后,将后验SNR加权能量差应用于去噪语音信号的扩展基音段以检测语音活动。我们使用两个数据库,RATS和Aurora-2,其中包含各种各样的噪声条件下,所提出的方法的VAD性能进行评估。在RedDots 2016挑战数据库及其噪声破坏版本上,进一步评估了rVAD方法的说话人验证性能。实验结果表明,rVAD是优于现有的方法。此外,我们提出了一个修改后的版本的rVAD计算密集的音高提取取代计算高效的频谱平坦度计算。修改后的版本显著降低了计算复杂度,但代价是VAD性能稍差,这在处理大量数据和在低资源设备上运行时是一个优势。rVAD的源代码是公开的。(C)2019爱思唯尔有限公司版权所有。
This paper presents an unsupervised segment-based method for robust voice activity detection (rVAD). The method consists of two passes of denoising followed by a voice activity detection (VAD) stage. In the first pass, high-energy segments in a speech signal are detected by using a posteriori signal-to-noise ratio (SNR) weighted energy difference and if no pitch is detected within a segment, the segment is considered as a high-energy noise segment and set to zero. In the second pass, the speech signal is denoised by a speech enhancement method, for which several methods are explored. Next, neighbouring frames with pitch are grouped together to form pitch segments, and based on speech statistics, the pitch segments are further extended from both ends in order to include both voiced and unvoiced sounds and likely non-speech parts as well. In the end, a posteriori SNR weighted energy difference is applied to the extended pitch segments of the denoised speech signal for detecting voice activity. We evaluate the VAD performance of the proposed method using two databases, RATS and Aurora-2, which contain a large variety of noise conditions. The rVAD method is further evaluated, in terms of speaker verification performance, on the RedDots 2016 challenge database and its noise-corrupted versions. Experiment results show that rVAD is compared favourably with a number of existing methods. In addition, we present a modified version of rVAD where computationally intensive pitch extraction is replaced by computationally efficient spectral flatness calculation. The modified version significantly reduces the computational complexity at the cost of moderately inferior VAD performance, which is an advantage when processing a large amount of data and running on low resource devices. The source code of rVAD is made publicly available. (C) 2019 Elsevier Ltd. All rights reserved.