Voice spoofing detector: A unified anti-spoofing framework
Voice spoofing detector: A unified anti-spoofing framework
复制标题
DOI:
10.1016/j.eswa.2022.116770
复制
发表时间:
2022-03-11
影响因子:
8.5
通讯作者:
Irtaza, Aun
中科院分区:
文献类型:
--
作者:
Javed, Ali;Malik, Khalid Mahmood;Irtaza, Aun
oice controlled systems (VCS) in Internet of Things (IoT), speaker verification systems, voice-based biometrics,and other voice-assistant-enabled systems are vulnerable to different spoofing attacks i.e., replay, cloning,cloned-replay, etc. VCS are not only susceptible to these attacks in a non-network environment, but theyare also vulnerable to multi-order spoofing attacks in networked IoT. Additionally, deepfakes with artificiallygenerated audio pose a great threat to the all systems having voice-interfaces. Most of the existing counter-measures against these voice spoofing attacks work for only one specific attack (e.g. voice replay) and fail togeneralize this for other classes of spoofing attacks. Additionally, generalization is also crucial for cross-corporaevaluation. Thus, there exists a need to develop a unified voice anti-spoofing framework capable of detectingmultiple spoofing attacks. This work presents a unified anti-spoofing framework that uses novel (ATCoP-GTCC)features to combat the variety of voice spoofing attacks. The proposed novel acoustic-ternary co-occurrencepatterns (ATCoP) encode the co-occurrence of similar patterns between the center and neighboring samples.Our experiments demonstrate that ATCoP can better capture the microphone induced distortions in replays,unnatural prosody and algorithmic artifacts in cloned samples, and both the distortions and artifacts in cloned-replays including compression on multi-hop attacks in the spoofing samples. The performance of ATCoP couldbe further enhanced by the Gammatone cepstral coefficients. To evaluate the effectiveness of the proposedanti-spoofing system for multi-order replay and cloned-replay attacks detection, we created a diverse voicespoofing detection corpus (VSDC) containing multi-order replay and cloned-replay audios against the bonafideand cloned audio recordings, respectively. Experimental results obtained on VSDC, ASVspoof 2019, Google'sLJ Speech, and YouTube deepfakes datasets illustrate the effectiveness of the proposed system in terms ofaccurate detection for a variety of voice spoofing attacks.