"Hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World

"Hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World
复制标题

DOI:
10.1145/3460120.3484742
复制
发表时间:
2021-09
期刊:
Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security
影响因子:
--
通讯作者:
Emily Wenger;Max Bronckers;Christian Cianfarani;Jenna Cryan;Angela Sha;Haitao Zheng;Ben Y. Zhao
Emily Wenger;Max Bronckers;Christian Cianfarani;Jenna Cryan;Angela Sha;Haitao Zheng;Ben Y. Zhao
中科院分区:
其他
文献类型:
--
作者:
Emily Wenger;Max Bronckers;Christian Cianfarani;Jenna Cryan;Angela Sha;Haitao Zheng;Ben Y. Zhao

文献摘要

被引文献

相似文献

深度学习的进步引入了新一波的语音合成工具,能够产生听起来像是目标说话者所说的音频。如果成功,这些工具落入坏人之手,将会对人类和软件系统(又名机器)发起一系列强大的攻击。本文记录了一项综合实验研究的成果和结果,该研究涉及基于深度学习的语音合成攻击对人类听众和机器(例如说话人识别和语音登录系统)的影响。我们发现人类和机器都可以被合成语音可靠地愚弄,并且现有的针对合成语音的防御措施不足。这些发现强调需要提高认识并针对人类和机器的合成语音开发新的保护措施。
Advances in deep learning have introduced a new wave of voice synthesis tools, capable of producing audio that sounds as if spoken by a target speaker. If successful, such tools in the wrong hands will enable a range of powerful attacks against both humans and software systems (aka machines). This paper documents efforts and findings from a comprehensive experimental study on the impact of deep-learning based speech synthesis attacks on both human listeners and machines such as speaker recognition and voice-signin systems. We find that both humans and machines can be reliably fooled by synthetic speech, and that existing defenses against synthesized speech fall short. These findings highlight the need to raise awareness and develop new protections against synthetic speech for both humans and machines.