Investigating self-supervised front ends for speech spoofing countermeasures

Investigating self-supervised front ends for speech spoofing countermeasures
复制标题

DOI:
10.21437/odyssey.2022-14
复制
发表时间:
2021-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Xin Wang;J. Yamagishi
Xin Wang;J. Yamagishi
中科院分区:
其他
文献类型:
--
作者:
Xin Wang;J. Yamagishi

文献摘要

相似文献

自监督语音模型是一个快速发展的研究课题,许多预训练模型已经发布并用于各种下游任务。对于语音反欺骗,大多数对策(CM)使用信号处理算法来提取声学特征进行分类。在这项研究中,我们使用预训练的自监督语音模型作为欺骗CM的前端。我们研究了与自监督前端相结合的不同后端架构,微调前端的有效性以及使用不同预训练自监督模型的性能。我们的研究结果表明,当一个良好的预训练前端在ASVspoof 2019逻辑访问(LA)训练集上使用基于浅层或深层神经网络的后端进行微调时,所产生的CM不仅在2019 LA测试集上获得了较低的EER分数,而且在ASVspoof 2015,2021 LA和2021 deepfake测试集上的表现明显优于基线。子带分析进一步表明,CM主要使用特定频带中的信息来区分测试集中的真实和欺骗性试验。
Self-supervised speech model is a rapid progressing research topic, and many pre-trained models have been released and used in various down stream tasks. For speech anti-spoofing, most countermeasures (CMs) use signal processing algorithms to extract acoustic features for classification. In this study, we use pre-trained self-supervised speech models as the front end of spoofing CMs. We investigated different back end architectures to be combined with the self-supervised front end, the effectiveness of fine-tuning the front end, and the performance of using different pre-trained self-supervised models. Our findings showed that, when a good pre-trained front end was fine-tuned with either a shallow or a deep neural network-based back end on the ASVspoof 2019 logical access (LA) training set, the resulting CM not only achieved a low EER score on the 2019 LA test set but also significantly outperformed the baseline on the ASVspoof 2015, 2021 LA, and 2021 deepfake test sets. A sub-band analysis further demonstrated that the CM mainly used the information in a specific frequency band to discriminate the bona fide and spoofed trials across the test sets.