Noise robust voice activity detection using joint phase and magnitude based feature enhancement

Noise robust voice activity detection using joint phase and magnitude based feature enhancement
复制标题

DOI:
10.1007/s12652-017-0482-8
复制
发表时间:
2017-04
影响因子:
--
通讯作者:
Khomdet Phapatanaburi;Longbiao Wang;Zeyan Oo;Weifeng Li;S. Nakagawa;M. Iwahashi
Khomdet Phapatanaburi;Longbiao Wang;Zeyan Oo;Weifeng Li;S. Nakagawa;M. Iwahashi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Khomdet Phapatanaburi;Longbiao Wang;Zeyan Oo;Weifeng Li;S. Nakagawa;M. Iwahashi

文献摘要

被引文献

相似文献

最近,基于深度神经网络(DNN)的特征增强已经被提出用于许多语音应用。DNN增强功能比原始功能实现了更高的性能。然而,在大多数常规DNN训练期间,相位信息被丢弃。在本文中,我们提出了一种基于DNN的联合基于相位和幅度的特征(JPMF)增强(JPMF with DNN)和一种基于噪声感知训练(NAT)-DNN的JPMF增强(JPMF with NAT-DNN),用于噪声鲁棒的语音活动检测(VAD)。此外,为了提高所提出的特征增强的性能,还应用了所提出的基于相位和基于幅度的特征的分数的组合。具体而言,梅尔频率倒谱系数(MFCC)和梅尔频率增量相位(MFDP)被用作幅度和相位特征。实验结果表明,所提出的特征增强明显优于传统的基于幅度的特征增强。与单独的基于幅度和相位的DNN语音增强相比,所提出的JPMF与NAT-DNN方法实现了最佳的相对相等错误率(EER)。此外,使用JPMF和NAT-DNN的增强MFCC和MFDP的组合得分进一步提高了VAD性能。
Recently, deep neural network (DNN)-based feature enhancement has been proposed for many speech applications. DNN-enhanced features have achieved higher performance than raw features. However, phase information is discarded during most conventional DNN training. In this paper, we propose a DNN-based joint phase- and magnitude -based feature (JPMF) enhancement (JPMF with DNN) and a noise-aware training (NAT)-DNN-based JPMF enhancement (JPMF with NAT-DNN) for noise-robust voice activity detection (VAD). Moreover, to improve the performance of the proposed feature enhancement, a combination of the scores of the proposed phase- and magnitude-based features is also applied. Specifically, mel-frequency cepstral coefficients (MFCCs) and the mel-frequency delta phase (MFDP) are used as magnitude and phase features. The experimental results show that the proposed feature enhancement significantly outperforms the conventional magnitude-based feature enhancement. The proposed JPMF with NAT-DNN method achieves the best relative equal error rate (EER), compared with individual magnitude- and phase-based DNN speech enhancement. Moreover, the combined score of the enhanced MFCC and MFDP using JPMF with NAT-DNN further improves the VAD performance.