Learning to Listen and Listening to Learn: Spoofed Audio Detection Through Linguistic Data Augmentation

Learning to Listen and Listening to Learn: Spoofed Audio Detection Through Linguistic Data Augmentation
复制标题

DOI:
10.1109/isi58743.2023.10297267
复制
发表时间:
2023-10
期刊:
2023 IEEE International Conference on Intelligence and Security Informatics (ISI)
影响因子:
--
通讯作者:
Zahra Khanjani;Lavon Davis;Anna Tuz;Kifekachukwu Nwosu;Christine Mallinson;V. P. Janeja
Zahra Khanjani;Lavon Davis;Anna Tuz;Kifekachukwu Nwosu;Christine Mallinson;V. P. Janeja
中科院分区:
其他
文献类型:
--
作者:
Zahra Khanjani;Lavon Davis;Anna Tuz;Kifekachukwu Nwosu;Christine Mallinson;V. P. Janeja

文献摘要

相似文献

无论是人类还是机器生成的欺骗音频,都会导致欺骗和虚假信息,因此是一个社会挑战。这项研究通过一种新的方法推进了对欺骗音频的检测,该方法通过结合语言信息来增加关于音频数据的知识。对于英语音频样本,社会语言学专家使用知觉方法来倾听音频提示,并根据英语口语中频繁出现的语音和音位特征,使用二进制标签来指示一组语音样本的感知真实性。然后将这些专家定义的语言特征(EDFL)用于有监督的欺骗音频检测方法,以增强AI模型。建立了一种基于音频数据本身和EDFL的多域特征的集成方法来评估欺骗音频检测,并演示了EDLF如何改进传统的欺骗音频检测方法。我们发现,在经过评估的单一模型和集成模型的训练和测试数据集上,使用专业形成的语言注释来扩充音频数据显著地提高了欺骗音频检测的准确性。我们的发现表明,使用感知语言技术来增加音频数据,作为一种人类识别的方法,以增强基于人工智能的欺骗音频检测方法,是一条很有前途的途径。这些功能还为新音频片段上的直接语言注释奠定了基础,以实现稳健的欺骗音频检测。
Spoofed audio, both human or machine generated, causes deception and disinformation and as such is a societal challenge. This study advances the detection of spoofed audio through a novel approach that augments knowledge about audio data by incorporating linguistic information. Using perceptual methods, for English audio samples, experts in sociolinguistics listened for audio cues, and used binary labels to indicate the perceived authenticity of a set of speech samples, based on phonetic and phonological features that occur frequently in spoken English. These Expert Defined Linguistic Features (EDLFs) were then used in supervised spoofed audio detection methods to augment AI models. An ensemble method based on multi-domain features both from the audio data itself and the EDLFs was also created to evaluate the spoofed audio detection, and to demonstrate how EDLFs can improve traditional methods of spoofed audio detection. We found that augmenting the audio data with expertinformed linguistic annotation increased the accuracy of spoofed audio detection significantly in both the training and testing datasets across the evaluated single and ensemble models. Our findings indicate the promising avenue of augmenting audio data with perceptual linguistic techniques, as a method of human discernment, to enhance AI-based approaches for spoofed audio detection. These features also establish a foundation for direct linguistic annotations on new audio clips for robust spoofed audio detection.