Voice spoofing detector: A unified anti-spoofing framework

Voice spoofing detector: A unified anti-spoofing framework
复制标题

DOI:
10.1016/j.eswa.2022.116770
复制
发表时间:
2022-03-11
影响因子:
8.5
通讯作者:
Irtaza, Aun
Irtaza, Aun
中科院分区:
计算机科学1区
文献类型:
--
作者:
Javed, Ali;Malik, Khalid Mahmood;Irtaza, Aun

文献摘要

被引文献

相似文献

物联网 (IoT) 中的语音控制系统 (VCS)、说话人验证系统、基于语音的生物识别系统和其他支持语音助手的系统容易受到不同的欺骗攻击,即重放、克隆、克隆重放等。VCS 不仅在非网络环境中容易受到这些攻击,而且在联网物联网中也容易受到多阶欺骗攻击。此外,人工生成音频的深度伪造对所有具有语音接口的系统构成了巨大威胁。针对这些语音欺骗攻击的大多数现有对策仅适用于一种特定攻击(例如语音重放),并且无法将其推广到其他类别的欺骗攻击。此外,泛化对于跨企业评估也至关重要。因此,需要开发一种能够检测多种欺骗攻击的统一语音反欺骗框架。这项工作提出了一个统一的反欺骗框架,该框架使用新颖的(ATCoP-GTCC)功能来对抗各种语音欺骗攻击。所提出的新颖的声学三元共现模式(ATCoP)对中心样本和相邻样本之间相似模式的共现进行编码。我们的实验表明,ATCoP 可以更好地捕获重放中麦克风引起的失真、克隆样本中的不自然韵律和算法伪影,以及克隆重放中的失真和伪影,包括欺骗样本中多跳攻击的压缩。 Gammatone 倒谱系数可以进一步增强 ATCoP 的性能。为了评估所提出的多阶重放和克隆重放攻击检测反欺骗系统的有效性,我们创建了一个多样化的语音欺骗检测语料库(VSDC),其中分别包含针对善意和克隆录音的多阶重放和克隆重放音频。在 VSDC、ASVspoof 2019、Google 的 LJ Speech 和 YouTube deepfakes 数据集上获得的实验结果说明了所提出的系统在准确检测各种语音欺骗攻击方面的有效性。
oice controlled systems (VCS) in Internet of Things (IoT), speaker verification systems, voice-based biometrics,and other voice-assistant-enabled systems are vulnerable to different spoofing attacks i.e., replay, cloning,cloned-replay, etc. VCS are not only susceptible to these attacks in a non-network environment, but theyare also vulnerable to multi-order spoofing attacks in networked IoT. Additionally, deepfakes with artificiallygenerated audio pose a great threat to the all systems having voice-interfaces. Most of the existing counter-measures against these voice spoofing attacks work for only one specific attack (e.g. voice replay) and fail togeneralize this for other classes of spoofing attacks. Additionally, generalization is also crucial for cross-corporaevaluation. Thus, there exists a need to develop a unified voice anti-spoofing framework capable of detectingmultiple spoofing attacks. This work presents a unified anti-spoofing framework that uses novel (ATCoP-GTCC)features to combat the variety of voice spoofing attacks. The proposed novel acoustic-ternary co-occurrencepatterns (ATCoP) encode the co-occurrence of similar patterns between the center and neighboring samples.Our experiments demonstrate that ATCoP can better capture the microphone induced distortions in replays,unnatural prosody and algorithmic artifacts in cloned samples, and both the distortions and artifacts in cloned-replays including compression on multi-hop attacks in the spoofing samples. The performance of ATCoP couldbe further enhanced by the Gammatone cepstral coefficients. To evaluate the effectiveness of the proposedanti-spoofing system for multi-order replay and cloned-replay attacks detection, we created a diverse voicespoofing detection corpus (VSDC) containing multi-order replay and cloned-replay audios against the bonafideand cloned audio recordings, respectively. Experimental results obtained on VSDC, ASVspoof 2019, Google'sLJ Speech, and YouTube deepfakes datasets illustrate the effectiveness of the proposed system in terms ofaccurate detection for a variety of voice spoofing attacks.