One-Class Learning Towards Synthetic Voice Spoofing Detection

One-Class Learning Towards Synthetic Voice Spoofing Detection
复制标题

合成语音欺骗检测的一堂课学习

DOI:
10.1109/lsp.2021.3076358
复制
发表时间:
2020-10
影响因子:
3.9
通讯作者:
You Zhang;Fei Jiang;Z. Duan
You Zhang;Fei Jiang;Z. Duan
中科院分区:
工程技术2区
文献类型:
--
作者:
You Zhang;Fei Jiang;Z. Duan

文献摘要

相似文献

人的声音可以用来验证说话人的身份,但自动说话人验证(ASV)系统容易受到语音欺骗攻击,如模仿,重放,文本到语音,和语音转换。最近,研究人员开发了反欺骗技术,以提高ASV系统对欺骗攻击的可靠性。然而,大多数方法在实际应用中遇到的困难,检测未知的攻击,往往有不同的统计分布从已知的攻击。特别是,合成语音欺骗算法的快速发展正在产生越来越强大的攻击,使ASV系统面临不可见的攻击风险。在这项工作中,我们提出了一个反欺骗系统来检测未知的合成语音欺骗攻击(即,文本到语音或语音转换)。其核心思想是压缩真实的语音表示,并在嵌入空间中注入一个角度余量来分离欺骗攻击。在不诉诸任何数据增强方法的情况下,我们提出的系统在ASVspoof 2019挑战逻辑访问场景的评估集上实现了2.19%的等错误率(EER),优于所有现有的单一系统(即,没有模型的集合)。
Human voices can be used to authenticate the identity of the speaker, but the automatic speaker verification (ASV) systems are vulnerable to voice spoofing attacks, such as impersonation, replay, text-to-speech, and voice conversion. Recently, researchers developed anti-spoofing techniques to improve the reliability of ASV systems against spoofing attacks. However, most methods encounter difficulties in detecting unknown attacks in practical use, which often have different statistical distributions from known attacks. Especially, the fast development of synthetic voice spoofing algorithms is generating increasingly powerful attacks, putting the ASV systems at risk of unseen attacks. In this work, we propose an anti-spoofing system to detect unknown synthetic voice spoofing attacks (i.e., text-to-speech or voice conversion) using one-class learning. The key idea is to compact the bona fide speech representation and inject an angular margin to separate the spoofing attacks in the embedding space. Without resorting to any data augmentation methods, our proposed system achieves an equal error rate (EER) of 2.19% on the evaluation set of ASVspoof 2019 Challenge logical access scenario, outperforming all existing single systems (i.e., those without model ensemble).