Constrained Optimization with Dynamic Bound-scaling for Effective NLPBackdoor Defense

Constrained Optimization with Dynamic Bound-scaling for Effective NLPBackdoor Defense
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
--
影响因子:
--
通讯作者:
Guangyu Shen;Yingqi Liu;Guanhong Tao;Qiuling Xu;Zhuo Zhang;Shengwei An;Shiqing Ma;X. Zhang
Guangyu Shen;Yingqi Liu;Guanhong Tao;Qiuling Xu;Zhuo Zhang;Shengwei An;Shiqing Ma;X. Zhang
中科院分区:
其他
文献类型:
--
作者:
Guangyu Shen;Yingqi Liu;Guanhong Tao;Qiuling Xu;Zhuo Zhang;Shengwei An;Shiqing Ma;X. Zhang

文献摘要

相似文献

我们开发了一种新的优化方法NLP后门反演。我们利用softmax函数中动态降低的温度系数来为优化器提供不断变化的损失景观,以便该过程逐渐关注地面实况触发器,该触发器表示为凸船体中的独热值。我们的方法还具有一个温度回滚机制,以远离局部最优值,利用观察到局部最优值可以在NLP触发器反演中很容易确定(而不是在一般优化中)。我们在1600多个模型上评估了该技术(其中大约一半注入了后门),包括3个流行的NLP任务,4种不同的后门攻击和7种架构。我们的研究结果表明,该技术能够有效地检测和删除后门,优于4种基线方法。
We develop a novel optimization method for NLPbackdoor inversion. We leverage a dynamically reducing temperature coefficient in the softmax function to provide changing loss landscapes to the optimizer such that the process gradually focuses on the ground truth trigger, which is denoted as a one-hot value in a convex hull. Our method also features a temperature rollback mechanism to step away from local optimals, exploiting the observation that local optimals can be easily deter-mined in NLP trigger inversion (while not in general optimization). We evaluate the technique on over 1600 models (with roughly half of them having injected backdoors) on 3 prevailing NLP tasks, with 4 different backdoor attacks and 7 architectures. Our results show that the technique is able to effectively and efficiently detect and remove backdoors, outperforming 4 baseline methods.