The Feasibility and Inevitability of Stealth Attacks

The Feasibility and Inevitability of Stealth Attacks
复制标题

DOI:
10.1093/imamat/hxad027
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
I. Tyukin;D. Higham;Eliyas Woldegeorgis;Alexander N Gorban
I. Tyukin;D. Higham;Eliyas Woldegeorgis;Alexander N Gorban
中科院分区:
其他
文献类型:
--
作者:
I. Tyukin;D. Higham;Eliyas Woldegeorgis;Alexander N Gorban

文献摘要

相似文献

我们开发和研究新的对抗性扰动,使攻击者能够控制包括深度学习神经网络在内的通用人工智能 (AI) 系统中的决策。与对抗性数据修改相反,我们这里考虑的攻击机制涉及对人工智能系统本身的改变。这种秘密攻击可能是由软件开发团队中顽皮、腐败或心怀不满的成员发起的。它也可以由那些希望利用“人工智能民主化”议程的人来实现,其中网络架构和经过训练的参数集是公开共享的。我们开发了一系列新的可实施的攻击策略并进行了分析,表明隐形攻击很有可能变得透明,即系统性能在攻击者未知的固定验证集上保持不变,同时在感兴趣的触发输入上引发任何所需的输出。攻击者只需要估计验证集的大小和人工智能相关潜在空间的分布。就深度学习神经网络而言,我们证明单神经元攻击是可能的——对与单个神经元相关的权重和偏差进行修改——揭示了过度参数化产生的漏洞。我们在两个标准图像数据集上使用最先进的架构来说明这些概念。在理论和计算结果的指导下,我们还提出了防范隐形攻击的策略。
We develop and study new adversarial perturbations that enable an attacker to gain control over decisions in generic Artificial Intelligence (AI) systems including deep learning neural networks. In contrast to adversarial data modification, the attack mechanism we consider here involves alterations to the AI system itself. Such a stealth attack could be conducted by a mischievous, corrupt or disgruntled member of a software development team. It could also be made by those wishing to exploit a “democratization of AI” agenda, where network architectures and trained parameter sets are shared publicly. We develop a range of new implementable attack strategies with accompanying analysis, showing that with high probability a stealth attack can be made transparent, in the sense that system performance is unchanged on a fixed validation set which is unknown to the attacker, while evoking any desired output on a trigger input of interest. The attacker only needs to have estimates of the size of the validation set and the spread of the AI’s relevant latent space. In the case of deep learning neural networks, we show that a one neuron attack is possible—a modification to the weights and bias associated with a single neuron—revealing a vulnerability arising from over-parameterization. We illustrate these concepts using state of the art architectures on two standard image data sets. Guided by the theory and computational results, we also propose strategies to guard against stealth attacks.