AEVA: Black-box Backdoor Detection Using Adversarial Extreme Value Analysis

AEVA: Black-box Backdoor Detection Using Adversarial Extreme Value Analysis
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Junfeng Guo;Ang Li;Cong Liu
Junfeng Guo;Ang Li;Cong Liu
中科院分区:
其他
文献类型:
--
作者:
Junfeng Guo;Ang Li;Cong Liu

文献摘要

相似文献

深度神经网络(DNN)被证明容易受到后门攻击。后门通常通过将后门触发器注入训练示例来嵌入目标DNN中,这可能导致目标DNN错误分类附加有后门触发器的输入。现有的后门检测方法通常需要访问原始中毒训练数据、目标DNN的参数或每个给定输入的预测置信度,这在许多现实世界的应用中是不切实际的,例如,设备上部署的DNN。我们解决了黑盒硬标签后门检测问题,其中DNN是完全黑盒的,只有它的最终输出标签是可访问的。我们从优化的角度来处理这个问题,并表明后门检测的目标是有界的敌对目标。进一步的理论和实证研究表明,这种对抗性目标会导致具有高度偏斜分布的解决方案;在后门感染示例的对抗性映射中经常观察到奇点,我们称之为对抗性奇点现象。基于这一观察,我们提出了对抗极值分析(AEVA)来检测黑盒神经网络中的后门。AEVA是基于对抗地图的极值分析,从蒙特-卡罗梯度估计计算。通过对多个常见任务和后门攻击的大量实验证明,我们的方法在黑盒硬标签场景下检测后门攻击是有效的。
Deep neural networks (DNNs) are proved to be vulnerable against backdoor attacks. A backdoor is often embedded in the target DNNs through injecting a backdoor trigger into training examples, which can cause the target DNNs misclassify an input attached with the backdoor trigger. Existing backdoor detection methods often require the access to the original poisoned training data, the parameters of the target DNNs, or the predictive confidence for each given input, which are impractical in many real-world applications, e.g., on-device deployed DNNs. We address the black-box hard-label backdoor detection problem where the DNN is fully black-box and only its final output label is accessible. We approach this problem from the optimization perspective and show that the objective of backdoor detection is bounded by an adversarial objective. Further theoretical and empirical studies reveal that this adversarial objective leads to a solution with highly skewed distribution; a singularity is often observed in the adversarial map of a backdoor-infected example, which we call the adversarial singularity phenomenon. Based on this observation, we propose the adversarial extreme value analysis(AEVA) to detect backdoors in black-box neural networks. AEVA is based on an extreme value analysis of the adversarial map, computed from the monte-carlo gradient estimation. Evidenced by extensive experiments across multiple popular tasks and backdoor attacks, our approach is shown effective in detecting backdoor attacks under the black-box hard-label scenarios.