Learning to Listen and Listening to Learn: Spoofed Audio Detection Through Linguistic Data Augmentation
Learning to Listen and Listening to Learn: Spoofed Audio Detection Through Linguistic Data Augmentation
复制标题
DOI:
10.1109/isi58743.2023.10297267
复制
发表时间:
2023-10
期刊:
影响因子:
--
通讯作者:
Zahra Khanjani;Lavon Davis;Anna Tuz;Kifekachukwu Nwosu;Christine Mallinson;V. P. Janeja
中科院分区:
文献类型:
--
作者:
Zahra Khanjani;Lavon Davis;Anna Tuz;Kifekachukwu Nwosu;Christine Mallinson;V. P. Janeja
Spoofed audio, both human or machine generated, causes deception and disinformation and as such is a societal challenge. This study advances the detection of spoofed audio through a novel approach that augments knowledge about audio data by incorporating linguistic information. Using perceptual methods, for English audio samples, experts in sociolinguistics listened for audio cues, and used binary labels to indicate the perceived authenticity of a set of speech samples, based on phonetic and phonological features that occur frequently in spoken English. These Expert Defined Linguistic Features (EDLFs) were then used in supervised spoofed audio detection methods to augment AI models. An ensemble method based on multi-domain features both from the audio data itself and the EDLFs was also created to evaluate the spoofed audio detection, and to demonstrate how EDLFs can improve traditional methods of spoofed audio detection. We found that augmenting the audio data with expertinformed linguistic annotation increased the accuracy of spoofed audio detection significantly in both the training and testing datasets across the evaluated single and ensemble models. Our findings indicate the promising avenue of augmenting audio data with perceptual linguistic techniques, as a method of human discernment, to enhance AI-based approaches for spoofed audio detection. These features also establish a foundation for direct linguistic annotations on new audio clips for robust spoofed audio detection.