DNN-Based Amplitude and Phase Feature Enhancement for Noise Robust Speaker Identification

DNN-Based Amplitude and Phase Feature Enhancement for Noise Robust Speaker Identification
复制标题

DOI:
10.21437/interspeech.2016-717
复制
发表时间:
2016-09
期刊:
--
影响因子:
--
通讯作者:
Zeyan Oo;Yuta Kawakami;Longbiao Wang;S. Nakagawa;Xiong Xiao;M. Iwahashi
Zeyan Oo;Yuta Kawakami;Longbiao Wang;S. Nakagawa;Xiong Xiao;M. Iwahashi
中科院分区:
其他
文献类型:
--
作者:
Zeyan Oo;Yuta Kawakami;Longbiao Wang;S. Nakagawa;Xiong Xiao;M. Iwahashi

文献摘要

被引文献

相似文献

语音信号相位信息的重要性日益受到关注。许多研究表明,幅度和相位特征的系统组合对于提高噪声环境下的说话人识别性能是有效的。另一方面,通常采用语音增强方法来减少噪声的影响。然而,这种方法仅增强了幅度谱,因此噪声相位谱用于重建估计信号。近年来,基于 DNN 的特征增强在鲁棒语音处理方面得到了深入研究。这种方法预计对于基于阶段的特征也有效。在本文中,我们提出使用深度神经网络(DNN)对幅度和相位特征进行特征空间增强以进行说话人识别。我们使用梅尔频率倒谱系数作为幅度特征,并使用修改的群延迟倒谱系数作为相位特征。基于幅度和相位的特征同时增强是有效的,与单独特征增强相比,相对误差降低了约24%。
The importance of the phase information of speech signal is gathering attention. Many researches indicate system combination of the amplitude and phase features is effective for improving speaker recognition performance under noisy environments. On the other hand, speech enhancement approach is taken usually to reduce the influence of noises. However, this approach only enhances the amplitude spectrum, therefor noisy phase spectrum is used for reconstructing the estimated signal. Recent years, DNN based feature enhancement is studied intensively for robust speech processing. This approach is expected to be effective also for phase-based feature. In this paper, we propose feature space enhancement of amplitude and phase features using deep neural network (DNN) for speaker identification. We used mel-frequency cepstral coefficients as an amplitude feature, and modified group delay cepstral coefficients as a phase feature. Simultaneous enhancement of amplitude and phase based feature was effective, and it achieved about 24% relative error reduction comparing with individual feature enhancement.