A Sweet Rabbit Hole by DARCY: Using Honeypots to Detect Universal Trigger’s Adversarial Attacks

A Sweet Rabbit Hole by DARCY: Using Honeypots to Detect Universal Trigger’s Adversarial Attacks
复制标题

DOI:
10.18653/v1/2021.acl-long.296
复制
发表时间:
2020-11
期刊:
--
影响因子:
--
通讯作者:
Thai Le;Noseong Park;Dongwon Lee
Thai Le;Noseong Park;Dongwon Lee
中科院分区:
其他
文献类型:
--
作者:
Thai Le;Noseong Park;Dongwon Lee

文献摘要

被引文献

相似文献

通用触发器(UniTrigger)是最近提出的一种强大的对抗性文本攻击方法。利用基于学习的机制,UniTrigger生成固定短语,当添加到任何良性输入时,可以将文本神经网络(NN)模型的预测精度降低到目标类的接近零。为了防御这种可能造成重大危害的攻击,本文借鉴了网络安全界的蜜罐概念,提出了一种基于蜜罐的UniTrigger防御框架Darcy。Darcy贪婪地搜索并向NN模型中注入多个陷门,以“诱饵和捕捉”潜在的攻击。通过对四个公共数据集的综合实验,我们表明,Darcy在大多数情况下以高达99%的TPR和低于2%的FPR检测到UniTrigger的对手攻击,同时将对干净输入的预测精度(在F1中)保持在1%的范围内。我们还展示了具有多个陷门的Darcy对于攻击者的不同知识和技能水平的不同攻击场景也是健壮的。我们在https://github.com/lethaiq/ACL2021-DARCY-HoneypotDefenseNLP.上发布达西的源代码
The Universal Trigger (UniTrigger) is a recently-proposed powerful adversarial textual attack method. Utilizing a learning-based mechanism, UniTrigger generates a fixed phrase that, when added to any benign inputs, can drop the prediction accuracy of a textual neural network (NN) model to near zero on a target class. To defend against this attack that can cause significant harm, in this paper, we borrow the “honeypot” concept from the cybersecurity community and propose DARCY, a honeypot-based defense framework against UniTrigger. DARCY greedily searches and injects multiple trapdoors into an NN model to “bait and catch” potential attacks. Through comprehensive experiments across four public datasets, we show that DARCY detects UniTrigger’s adversarial attacks with up to 99% TPR and less than 2% FPR in most cases, while maintaining the prediction accuracy (in F1) for clean inputs within a 1% margin. We also demonstrate that DARCY with multiple trapdoors is also robust to a diverse set of attack scenarios with attackers’ varying levels of knowledge and skills. We release the source code of DARCY at: https://github.com/lethaiq/ACL2021-DARCY-HoneypotDefenseNLP.