Data Augmentation Using McAdams-Coefficient-Based Speaker Anonymization for Fake Audio Detection

Data Augmentation Using McAdams-Coefficient-Based Speaker Anonymization for Fake Audio Detection
复制标题

DOI:
10.21437/interspeech.2022-10088
复制
发表时间:
2022-09
期刊:
--
影响因子:
--
通讯作者:
Kai Li;Sheng Li;Xugang Lu;M. Akagi;Meng Liu-;Lin Zhang;Chang Zeng;Longbiao Wang;J. Dang;M. Unoki
Kai Li;Sheng Li;Xugang Lu;M. Akagi;Meng Liu-;Lin Zhang;Chang Zeng;Longbiao Wang;J. Dang;M. Unoki
中科院分区:
其他
文献类型:
--
作者:
Kai Li;Sheng Li;Xugang Lu;M. Akagi;Meng Liu-;Lin Zhang;Chang Zeng;Longbiao Wang;J. Dang;M. Unoki

文献摘要

相似文献

假音频检测(FAD)是一种区分人工语音和自然语音的技术。在大多数FAD系统中,从声学语音中去除不相关的特征,同时只保留鲁棒的判别特征是必不可少的。直观地说,在FAD任务中,应该抑制声语音中纠缠的说话人信息。特别是在基于深度神经网络(DNN)的FAD系统中,学习系统可能会从训练数据集中学习说话人信息,而不能很好地泛化测试数据集。在本文中,我们提出使用说话人匿名化(SA)技术在将声学语音输入到基于dnn的FAD系统之前抑制说话人信息。我们采用了McAdams-coefficient-based SA (MC-SA)算法,期望在基于dnn的FAD学习中不会涉及到纠缠的说话人信息。基于这一思路,我们实现了一个基于光卷积神经网络双向长短期记忆(LCNN-BLSTM)的FAD系统,并在音频深度合成检测挑战(ADD2022)数据集上进行了实验。结果表明,从声学语音中去除说话人信息提高了ADD2022在第一道道上的相对性能
Fake audio detection (FAD) is a technique to distinguish synthetic speech from natural speech. In most FAD systems, removing irrelevant features from acoustic speech while keeping only robust discriminative features is essential. Intuitively, speaker information entangled in acoustic speech should be suppressed for the FAD task. Particularly in a deep neural network (DNN)-based FAD system, the learning system may learn speaker information from a training dataset and cannot generalize well on a testing dataset. In this paper, we propose to use the speaker anonymization (SA) technique to suppress speaker information from acoustic speech before inputting it into a DNN-based FAD system. We adopted the McAdams-coefficient-based SA (MC-SA) algorithm, and this is expected that the entangled speaker information will not be involved in the DNN-based FAD learning. Based on this idea, we implemented a light convolutional neural network bidirectional long short-term memory (LCNN-BLSTM)-based FAD system and conducted experiments on the Audio Deep Synthesis Detection Challenge (ADD2022) datasets. The results showed that removing the speaker information from acoustic speech improved the relative performance in the first track of ADD2022