Robust Speaker Recognition Based on Single-Channel and Multi-Channel Speech Enhancement

Robust Speaker Recognition Based on Single-Channel and Multi-Channel Speech Enhancement
复制标题

DOI:
10.1109/taslp.2020.2986896
复制
发表时间:
2020
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
H. Taherian;Zhong-Qiu Wang;Jorge Chang;Deliang Wang
H. Taherian;Zhong-Qiu Wang;Jorge Chang;Deliang Wang
中科院分区:
其他
文献类型:
--
作者:
H. Taherian;Zhong-Qiu Wang;Jorge Chang;Deliang Wang

文献摘要

相似文献

用于说话人识别的深度神经网络(DNN)嵌入最近引起了广泛的关注。与i向量相比,它们对噪声和房间混响更鲁棒,因为DNN利用了大规模训练。本文讨论了当DNN嵌入用于说话人识别时,语音增强方法是否仍然有用的问题。我们研究了在强扩散噪声和混响同时存在的条件下,基于x向量的文本无关说话人确认的单通道和多通道语音增强。单声道(单声道)语音增强是基于复杂的频谱映射,并适用于个别麦克风。我们使用基于掩蔽的最小方差无失真响应(MVDR)波束形成器及其秩1近似的多通道语音增强。我们提出了一种新的方法,从估计的复谱图推导出时频掩模。此外,我们研究伽玛频率倒谱系数(GFCCs)作为强大的扬声器功能。在NIST SRE 2010重传语料库上的系统评估和比较表明,单声道和多声道语音增强的性能都明显优于x-向量的性能,并且我们的协方差矩阵估计对于MVDR波束形成器是有效的。
Deep neural network (DNN) embeddings for speaker recognition have recently attracted much attention. Compared to i-vectors, they are more robust to noise and room reverberation as DNNs leverage large-scale training. This article addresses the question of whether speech enhancement approaches are still useful when DNN embeddings are used for speaker recognition. We investigate single- and multi-channel speech enhancement for text-independent speaker verification based on x-vectors in conditions where strong diffuse noise and reverberation are both present. Single-channel (monaural) speech enhancement is based on complex spectral mapping and is applied to individual microphones. We use masking-based minimum variance distortion-less response (MVDR) beamformer and its rank-1 approximation for multi-channel speech enhancement. We propose a novel method of deriving time-frequency masks from the estimated complex spectrogram. In addition, we investigate gammatone frequency cepstral coefficients (GFCCs) as robust speaker features. Systematic evaluations and comparisons on the NIST SRE 2010 retransmitted corpus show that both monaural and multi-channel speech enhancement significantly outperform x-vector's performance, and our covariance matrix estimate is effective for the MVDR beamformer.