LoneNeuron: A Highly-Effective Feature-Domain Neural Trojan Using Invisible and Polymorphic Watermarks

LoneNeuron: A Highly-Effective Feature-Domain Neural Trojan Using Invisible and Polymorphic Watermarks
复制标题

DOI:
10.1145/3548606.3560678
复制
发表时间:
2022-11
期刊:
Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security
影响因子:
--
通讯作者:
Zeyan Liu;Fengjun Li;Zhu Li;Bo Luo
Zeyan Liu;Fengjun Li;Zhu Li;Bo Luo
中科院分区:
其他
文献类型:
--
作者:
Zeyan Liu;Fengjun Li;Zhu Li;Bo Luo

文献摘要

相似文献

深度神经网络(DNN)在实际应用中的广泛采用引发了越来越多的安全担忧。嵌入到预先训练的神经网络中的神经木马程序是对DNN模型供应链的有害攻击。当某些隐形触发器出现在输入中时,它们会生成错误输出。虽然数据中毒攻击已经在文献中得到了很好的研究,但代码中毒和模型中毒后门直到最近才开始引起人们的注意。我们提出了一种新的模型中毒神经特洛伊木马,即LoneNeuron,它响应特征域模式转换为不可见的、样本特定的和多态的像素域水印。LoneNeuron具有较高的攻击特异性,在不影响主任务性能的情况下,实现100%的攻击成功率。利用LoneNeuron独特的水印多态特性,将同一特征域触发器解析为像素域中的多个水印,进一步提高了水印的随机性、隐蔽性和抗木马检测能力。广泛的实验表明,LoneNeuron可以逃脱最先进的特洛伊木马探测器。LoneNeuron~也是针对视觉转换器(VITS)的第一个有效的后门攻击。
The wide adoption of deep neural networks (DNNs) in real-world applications raises increasing security concerns. Neural Trojans embedded in pre-trained neural networks are a harmful attack against the DNN model supply chain. They generate false outputs when certain stealthy triggers appear in the inputs. While data-poisoning attacks have been well studied in the literature, code-poisoning and model-poisoning backdoors only start to attract attention until recently. We present a novel model-poisoning neural Trojan, namely LoneNeuron, which responds to feature-domain patterns that transform into invisible, sample-specific, and polymorphic pixel-domain watermarks. With high attack specificity, LoneNeuron achieves a 100% attack success rate, while not affecting the main task performance. With LoneNeuron's unique watermark polymorphism property, the same feature-domain trigger is resolved to multiple watermarks in the pixel domain, which further improves watermark randomness, stealthiness, and resistance against Trojan detection. Extensive experiments show that LoneNeuron could escape state-of-the-art Trojan detectors. LoneNeuron~is also the first effective backdoor attack against vision transformers (ViTs).