Spoofing Speech Detection Using Modified Relative Phase Information

Spoofing Speech Detection Using Modified Relative Phase Information
复制标题

DOI:
10.1109/jstsp.2017.2694139
复制
发表时间:
2017-04
影响因子:
7.5
通讯作者:
Longbiao Wang;S. Nakagawa;Zhaofeng Zhang;Yohei Yoshida;Yuta Kawakami
Longbiao Wang;S. Nakagawa;Zhaofeng Zhang;Yohei Yoshida;Yuta Kawakami
中科院分区:
工程技术1区
文献类型:
--
作者:
Longbiao Wang;S. Nakagawa;Zhaofeng Zhang;Yohei Yoshida;Yuta Kawakami

文献摘要

被引文献

相似文献

人类和欺骗(合成或转换)语音的检测已开始受到越来越多的关注。在本文中,提出了从傅立叶频谱中提取的修改相对相位(MRP)信息用于欺骗语音检测。由于使用当前合成或转换技术的欺骗语音中原始相位信息几乎完全丢失,因此一些相位信息提取方法,例如改进的群延迟特征和余弦相位特征,已被证明对于检测人类语音和欺骗语音是有效的。然而,现有的基于相位信息的特征无法获得非常高的欺骗语音检测性能,因为它们无法从语音中提取精确的相位信息。相对相位(RP)信息可以精确提取相位信息,已被证明对于说话人识别非常有效。本文将RP信息应用于欺骗语音检测,有望取得更好的欺骗检测性能。此外,本文还提出了两种对原始RP的改进处理技术,即伪基音同步和基于线性判别分析的全频带RP提取。在本研究中,MRP 信息还与梅尔频率倒谱系数 (MFCC) 和修改的群延迟相结合。使用 ASVspoof 2015:自动说话人验证欺骗和对策挑战数据集对所提出的方法进行了评估。结果表明,所提出的 MRP 信息显着优于 MFCC、改进的群延迟和其他基于相位信息的特征。对于开发数据集,等错误率 (EER) 从 MFCC 的 1.883%、修改后的群延迟的 0.567% 降低到 MRP 的 0.013%。通过将 RP 与 MFCC 和修改后的群延迟相结合,EER 降低至 0.003%。对于评估数据集,除了S10欺骗语音之外,MRP比基于幅度的特征和其他基于相位的特征获得了更好的性能。
The detection of human and spoofing (synthetic or converted) speech has started to receive an increasing amount of attention. In this paper, modified relative phase (MRP) information extracted from a Fourier spectrum is proposed for spoofing speech detection. Because original phase information is almost entirely lost in spoofing speech using current synthesis or conversion techniques, some phase information extraction methods, such as the modified group delay feature and cosine phase feature, have been shown to be effective for detecting human speech and spoofing speech. However, existing phase information-based features cannot obtain very high spoofing speech detection performance because they cannot extract precise phase information from speech. Relative phase (RP) information, which extracts phase information precisely, has been shown to be effective for speaker recognition. In this paper, RP information is applied to spoofing speech detection, and it is expected to achieve better spoofing detection performance. Furthermore, two modified processing techniques of the original RP, that is, pseudo pitch synchronization and linear discriminant analysis based full-band RP extraction, are proposed in this paper. In this study, MRP information is also combined with the Mel-frequency cepstral coefficient (MFCC) and modified group delay. The proposed method was evaluated using the ASVspoof 2015: Automatic Speaker Verification Spoofing and Countermeasures Challenge dataset. The results show that the proposed MRP information significantly outperforms the MFCC, modified group delay, and other phase information based features. For the development dataset, the equal error rate (EER) was reduced from 1.883% of the MFCC, 0.567% of the modified group delay to 0.013% of the MRP. By combining the RP with the MFCC and modified group delay, the EER was reduced to 0.003%. For the evaluation dataset, the MRP obtained much better performance than the magnitude-based feature and other phase-based features, except for S10 spoofing speech.