Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation

Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation
复制标题

DOI:
10.21437/odyssey.2022-16
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Hemlata Tak;M. Todisco;Xin Wang;Jee-weon Jung;J. Yamagishi;N. Evans
Hemlata Tak;M. Todisco;Xin Wang;Jee-weon Jung;J. Yamagishi;N. Evans
中科院分区:
其他
文献类型:
--
作者:
Hemlata Tak;M. Todisco;Xin Wang;Jee-weon Jung;J. Yamagishi;N. Evans

文献摘要

相似文献

电子欺骗对抗系统的性能从根本上取决于充分代表性的训练数据的使用。由于这通常是有限的,当前的解决方案通常缺乏对在野外遇到的攻击的概括。因此,需要制定策略来提高面对不受控制、不可预测的攻击时的可靠性。我们在本文中报告了我们以wav2vec 2.0前端的形式使用自监督学习的努力。尽管仅使用真实数据而非欺骗数据来学习初始基础表示,但我们获得了ASVspoof 2021 Logical Access和Deepfake数据库文献中报告的最低等错误率。当与数据增强相结合时,这些结果对应于相对于我们的基线系统的近90%的改进。
The performance of spoofing countermeasure systems depends fundamentally upon the use of sufficiently representative training data. With this usually being limited, current solutions typically lack generalisation to attacks encountered in the wild. Strategies to improve reliability in the face of uncontrolled, unpredictable attacks are hence needed. We report in this paper our efforts to use self-supervised learning in the form of a wav2vec 2.0 front-end with fine tuning. Despite initial base representations being learned using only bona fide data and no spoofed data, we obtain the lowest equal error rates reported in the literature for both the ASVspoof 2021 Logical Access and Deepfake databases. When combined with data augmentation,these results correspond to an improvement of almost 90% relative to our baseline system.