ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation

ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation
复制标题

DOI:
10.1145/3319535.3363216
复制
发表时间:
2019-11
期刊:
Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security
影响因子:
--
通讯作者:
Yingqi Liu;Wen-Chuan Lee;Guanhong Tao;Shiqing Ma;Yousra Aafer;X. Zhang
Yingqi Liu;Wen-Chuan Lee;Guanhong Tao;Shiqing Ma;Yousra Aafer;X. Zhang
中科院分区:
其他
文献类型:
--
作者:
Yingqi Liu;Wen-Chuan Lee;Guanhong Tao;Shiqing Ma;Yousra Aafer;X. Zhang

文献摘要

被引文献

相似文献

本文介绍了一种基于神经网络的AI模型来确定它们是否被转换的技术。预训练的AI模型可能包含通过训练或转化内部神经元重量注入的后门。当提供常规输入时,这些Trojaned模型正常运行,当输入用某些特殊模式(称为Trojan Trigger)盖上输入时,将其错误分类为特定的输出标签。我们开发了一种新型技术,该技术通过确定输出激活如何在神经元引入不同水平的刺激时如何改变内部神经元行为。无论提供的输入如何,都可能会损害特定输出标签的激活的神经元。然后,使用刺激分析结果通过优化程序对Trojan触发进行反向工程,以确认神经元确实受到损害。我们在177个特洛伊木马模型上评估了系统ABS,这些模型均采用各种攻击方法,这些攻击方法既针对输入空间和特征空间,并且具有各种木马的触发尺寸和形状,以及144个良性模型,这些模型接受了不同的数据和初始重量的培训值。这些模型属于7种不同的模型结构和6个不同的数据集,包括一些复杂的模型,例如ImageNet,VGG-FACE和RESNET110。我们的结果表明,当每个输出标签仅提供一个输入样本时,ABS非常有效,在大多数情况下(和许多100%)可以达到超过90%的检测率。它实际上超过了最先进的技术神经清洁技术,该技术需要大量输入样本和小型特洛伊木马触发器才能实现良好的性能。
This paper presents a technique to scan neural network based AI models to determine if they are trojaned. Pre-trained AI models may contain back-doors that are injected through training or by transforming inner neuron weights. These trojaned models operate normally when regular inputs are provided, and mis-classify to a specific output label when the input is stamped with some special pattern called trojan trigger. We develop a novel technique that analyzes inner neuron behaviors by determining how output activations change when we introduce different levels of stimulation to a neuron. The neurons that substantially elevate the activation of a particular output label regardless of the provided input is considered potentially compromised. Trojan trigger is then reverse-engineered through an optimization procedure using the stimulation analysis results, to confirm that a neuron is truly compromised. We evaluate our system ABS on 177 trojaned models that are trojaned with various attack methods that target both the input space and the feature space, and have various trojan trigger sizes and shapes, together with 144 benign models that are trained with different data and initial weight values. These models belong to 7 different model structures and 6 different datasets, including some complex ones such as ImageNet, VGG-Face and ResNet110. Our results show that ABS is highly effective, can achieve over 90% detection rate for most cases (and many 100%), when only one input sample is provided for each output label. It substantially out-performs the state-of-the-art technique Neural Cleanse that requires a lot of input samples and small trojan triggers to achieve good performance.