Speaker Recognition by Combining MFCC and Phase Information in Noisy Conditions

Speaker Recognition by Combining MFCC and Phase Information in Noisy Conditions
复制标题

DOI:
10.1587/transinf.e93.d.2397
复制
发表时间:
2010-09
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Longbiao Wang;Kazue Minami;Kazumasa Yamamoto;S. Nakagawa
Longbiao Wang;Kazue Minami;Kazumasa Yamamoto;S. Nakagawa
中科院分区:
其他
文献类型:
--
作者:
Longbiao Wang;Kazue Minami;Kazumasa Yamamoto;S. Nakagawa

文献摘要

被引文献

相似文献

本文研究了在噪声条件下相位对说话人识别的有效性,并将相位信息与梅尔倒谱系数(MFCC)联合收割机相结合。迄今为止,即使在噪声条件下,几乎所有的说话人识别方法都是基于MFCC的。对于主要捕获声道信息的MFCC,仅使用时域语音帧的傅里叶变换的幅度,而忽略相位信息。因为相位信息包括丰富的语音源信息,所以期望相位信息和MFCC的高互补性。此外,有研究表明,基于相位的特征对噪声具有很强的鲁棒性。在我们以前的研究中,提出了一种相位信息提取方法,该方法根据输入语音的削波位置对相位的变化进行归一化,并且相位信息和MFCC的组合的性能明显优于MFCC。在本文中,我们评估的鲁棒性建议的相位信息说话人识别在嘈杂的条件下。利用谱减法和带噪语音训练模型,分析了噪声环境下相位信息和MFCC对语音识别的影响。NTT数据库和JNAS(日本报纸文章句子)数据库中添加了平稳/非平稳噪声来评估我们提出的方法。MFCC优于干净语音的相位信息。另一方面,相位信息的退化明显小于噪声语音的MFCC。在许多情况下,通过干净的语音训练模型,相位信息的个体结果甚至优于MFCC。通过删除不可靠的帧(具有低能量/SN的帧),说话人识别性能得到显着改善。通过将相位信息与MFCC相结合,与基于标准MFCC的方法相比,说话人识别错误率降低了30%-60%。
In this paper, we investigate the effectiveness of phase for speaker recognition in noisy conditions and combine the phase information with mel-frequency cepstral coefficients (MFCCs). To date, almost speaker recognition methods are based on MFCCs even in noisy conditions. For MFCCs which dominantly capture vocal tract information, only the magnitude of the Fourier Transform of time-domain speech frames is used and phase information has been ignored. High complement of the phase information and MFCCs is expected because the phase information includes rich voice source information. Furthermore, some researches have reported that phase based feature was robust to noise. In our previous study, a phase information extraction method that normalizes the change variation in the phase depending on the clipping position of the input speech was proposed, and the performance of the combination of the phase information and MFCCs was remarkably better than that of MFCCs. In this paper, we evaluate the robustness of the proposed phase information for speaker identification in noisy conditions. Spectral subtraction, a method skipping frames with low energy/Signal-to-Noise (SN) and noisy speech training models are used to analyze the effect of the phase information and MFCCs in noisy conditions. The NTT database and the JNAS (Japanese Newspaper Article Sentences) database added with stationary/non-stationary noise were used to evaluate our proposed method. MFCCs outperformed the phase information for clean speech. On the other hand, the degradation of the phase information was significantly smaller than that of MFCCs for noisy speech. The individual result of the phase information was even better than that of MFCCs in many cases by clean speech training models. By deleting unreliable frames (frames having low energy/SN), the speaker identification performance was improved significantly. By integrating the phase information with MFCCs, the speaker identification error reduction rate was about 30%-60% compared with the standard MFCC-based method.