One-pixel Signature: Characterizing CNN Models for Backdoor Detection

One-pixel Signature: Characterizing CNN Models for Backdoor Detection
复制标题

DOI:
10.1007/978-3-030-58583-9_20
复制
发表时间:
2020-08
期刊:
--
影响因子:
--
通讯作者:
Shanjiaoyang Huang;Weiqi Peng;Zhiwei Jia;Z. Tu
Shanjiaoyang Huang;Weiqi Peng;Zhiwei Jia;Z. Tu
中科院分区:
其他
文献类型:
--
作者:
Shanjiaoyang Huang;Weiqi Peng;Zhiwei Jia;Z. Tu

文献摘要

相似文献

我们通过提出一种称为单像素签名的新表示来解决卷积神经网络(CNN)后门检测问题。我们的任务是检测/分类CNN模型是否被恶意插入了未知的特洛伊木马触发器。我们设计了一个像素的签名表示,以揭示干净和后门CNN模型的特征。在这里,每个CNN模型都与一个签名相关联,该签名是通过逐像素生成对抗值来创建的,该对抗值是类预测的最大变化的结果。单像素签名与CNN架构的设计选择以及它们是如何训练的无关。它可以有效地计算黑盒CNN模型,而无需访问网络参数。我们提出的单像素签名证明了比现有的竞争方法有实质性的改进(绝对检测准确率约为30%),用于后门CNN检测/分类。单像素签名是一种通用表示,可用于表征CNN模型,而不是后门检测。
We tackle the convolution neural networks (CNNs) backdoor detection problem by proposing a new representation called one-pixel signature. Our task is to detect/classify if a CNN model has been maliciously inserted with an unknown Trojan trigger or not. We design the one-pixel signature representation to reveal the characteristics of both clean and backdoored CNN models. Here, each CNN model is associated with a signature that is created by generating, pixel-by-pixel, an adversarial value that is the result of the largest change to the class prediction. The one-pixel signature is agnostic to the design choice of CNN architectures, and how they were trained. It can be computed efficiently for a black-box CNN model without accessing the network parameters. Our proposed one-pixel signature demonstrates a substantial improvement (by around 30% in the absolute detection accuracy) over the existing competing methods for backdoored CNN detection/classification. One-pixel signature is a general representation that can be used to characterize CNN models beyond backdoor detection.