Transforming acoustic characteristics to deceive playback spoofing countermeasures of speaker verification systems

Transforming acoustic characteristics to deceive playback spoofing countermeasures of speaker verification systems
复制标题

DOI:
10.1109/wifs.2018.8630764
复制
发表时间:
2018-09
期刊:
2018 IEEE International Workshop on Information Forensics and Security (WIFS)
影响因子:
--
通讯作者:
Fuming Fang;J. Yamagishi;I. Echizen;Md. Sahidullah;T. Kinnunen
Fuming Fang;J. Yamagishi;I. Echizen;Md. Sahidullah;T. Kinnunen
中科院分区:
其他
文献类型:
--
作者:
Fuming Fang;J. Yamagishi;I. Echizen;Md. Sahidullah;T. Kinnunen

文献摘要

相似文献

ASV (Automatic speaker verification)系统通过放音检测器来过滤放音攻击,保证验证的可靠性。由于目前的回放检测模型几乎总是使用真实语音和回放语音进行训练,因此有可能通过将播放语音的声学特性转换为接近真实语音的声学特性来降低其性能。一种方法是在重放之前加强从目标说话者那里“偷来的”言语。我们通过使用语音增强生成对抗网络来转换声学特征,测试了使用该方法进行重放攻击的有效性。实验结果表明,使用这种“增强窃取语音”方法可以显著提高ASVspoof 2017挑战中使用的基线和基于轻度卷积神经网络的方法的错误率。结果还表明,它的使用降低了基于高斯混合模型的通用背景模型的ASV系统的性能。因此,这种类型的攻击是一个迫切需要解决的问题。
Automatic speaker verification (ASV) systems use a playback detector to filter out playback attacks and ensure verification reliability. Since current playback detection models are almost always trained using genuine and playedback speech, it may be possible to degrade their performance by transforming the acoustic characteristics of the played-back speech close to that of the genuine speech. One way to do this is to enhance speech “stolen” from the target speaker before playback. We tested the effectiveness of a playback attack using this method by using the speech enhancement generative adversarial network to transform acoustic characteristics. Experimental results showed that use of this “enhanced stolen speech” method significantly increases the equal error rates for the baseline used in the ASVspoof 2017 challenge and for a light convolutional neural network-based method. The results also showed that its use degrades the performance of a Gaussian mixture modeluniversal background model-based ASV system. This type of attack is thus an urgent problem needing to be solved.