Auto Annotation of Linguistic Features for Audio Deepfake Discernment

Auto Annotation of Linguistic Features for Audio Deepfake Discernment
复制标题

DOI:
10.1609/aaaiss.v2i1.27682
复制
发表时间:
2024-01
期刊:
Proceedings of the AAAI Symposium Series
影响因子:
--
通讯作者:
Kifekachukwu Nwosu;Chloe Evered;Zahra Khanjani;Noshaba Bhalli;Lavon Davis;Christine Mallinson;V. P. Janeja
Kifekachukwu Nwosu;Chloe Evered;Zahra Khanjani;Noshaba Bhalli;Lavon Davis;Christine Mallinson;V. P. Janeja
中科院分区:
其他
文献类型:
--
作者:
Kifekachukwu Nwosu;Chloe Evered;Zahra Khanjani;Noshaba Bhalli;Lavon Davis;Christine Mallinson;V. P. Janeja

文献摘要

相似文献

我们提出了一种创新的方法来自动注释专家定义的语言特征(EDLF)作为音频时间序列中的连续性,以提高音频deepfake识别。在我们之前的工作中,这些语言特征-即音高,停顿,呼吸,辅音释放爆发和整体音频质量,由专家对整个音频信号进行标记-已被证明可以提高AI算法对音频deepfake的检测。现在,我们扩展我们的方法,以试验一种方法来自动注释对应于每个EDLF的时间序列中的连续性。我们开发了一个不一致的集合,即时间序列中的异常,使用跨多个不一致长度的矩阵配置文件来识别多种类型的EDLF。与语言专家密切合作,我们评估了音频信号数据中与EDLF重叠的不和谐音。我们的集成方法来检测跨多个不和谐长度的不和谐,实现了更高的准确性比使用单独的不和谐长度来检测EDLF。通过这种方法和领域验证,我们确定了使用时间序列序列捕获EDLF以补充领域专家的注释的可行性,以改进音频深度伪造检测。
We present an innovative approach to auto-annotate Expert Defined Linguistic Features (EDLFs) as subsequences in audio time series to improve audio deepfake discernment. In our prior work, these linguistic features – namely pitch, pause, breath, consonant release bursts, and overall audio quality, labeled by experts on the entire audio signal – have been shown to improve detection of audio deepfakes with AI algorithms. We now expand our approach to pilot a way to auto annotate subsequences in the time series that correspond to each EDLF. We developed an ensemble of discords, i.e. anomalies in time series, detected using matrix profiles across multiple discord lengths to identify multiple types of EDLFs. Working closely with linguistic experts, we evaluated where discords overlapped with EDLFs in the audio signal data. Our ensemble method to detect discords across multiple discord lengths achieves much higher accuracy than using individual discord lengths to detect EDLFs. With this approach and domain validation we establish the feasibility of using time series subsequences to capture EDLFs to supplement annotation by domain experts, for improved audio deepfake detection.