Detection of AI-Synthesized Speech Using Cepstral & Bispectral Statistics

Detection of AI-Synthesized Speech Using Cepstral & Bispectral Statistics
复制标题

使用倒谱检测人工智能合成语音

DOI:
--
复制
发表时间:
2020
期刊:
Conference on Multimedia Information Processing and Retrieval
影响因子:
--
通讯作者:
Priyanka Singh
Priyanka Singh
中科院分区:
--
文献类型:
--
作者:
A. Singh;Priyanka Singh

文献摘要

被引文献

相似文献

数字技术使难以想象的应用成为现实。拥有一些可以轻松编辑和操作的工具似乎令人兴奋,但它也引起了令人担忧的担忧,这些担忧可能会以语音克隆、重复或深度伪造的形式传播。验证语音的真实性是数字音频取证的主要问题之一。我们提出了一种利用双谱和倒谱分析来区分人类语音和人工智能合成语音的方法。与合成语音相比,高阶统计量与人类语音的相关性较小。此外,倒谱分析揭示了人类语音中的持久功率分量,而合成语音则缺少该分量。我们整合了这些分析,并提出了一个检测人工智能合成语音的模型。
Digital technology has made possible unimaginable applications come true. It seems exciting to have a handful of tools for easy editing and manipulation, but it raises alarming concerns that can propagate as speech clones, duplicates, or maybe deep fakes. Validating the authenticity of a speech is one of the primary problems of digital audio forensics. We propose an approach to distinguish human speech from AI synthesized speech exploiting the Bi-spectral and Cepstral analysis. Higher-order statistics have less correlation for human speech in comparison to a synthesized speech. Also, Cepstral analysis revealed a durable power component in human speech that is missing for a synthesized speech. We integrate both these analyses and propose a model to detect AI synthesized speech.