Defense against Synonym Substitution-based Adversarial Attacks via Dirichlet Neighborhood Ensemble

Defense against Synonym Substitution-based Adversarial Attacks via Dirichlet Neighborhood Ensemble
复制标题

DOI:
10.18653/v1/2021.acl-long.426
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Yi Zhou;Xiaoqing Zheng;Cho-Jui Hsieh;Kai-Wei Chang;Xuanjing Huang
Yi Zhou;Xiaoqing Zheng;Cho-Jui Hsieh;Kai-Wei Chang;Xuanjing Huang
中科院分区:
其他
文献类型:
--
作者:
Yi Zhou;Xiaoqing Zheng;Cho-Jui Hsieh;Kai-Wei Chang;Xuanjing Huang

文献摘要

被引文献

相似文献

尽管深度神经网络在许多NLP任务中取得了突出的性能,但它们很容易受到对手例子的影响。我们提出了Dirichlet邻域集成(DNE),这是一种随机方法,用于训练稳健的模型来防御基于同义词替换的攻击。在训练过程中,DNE通过从单词及其同义词所覆盖的凸壳中采样输入句子中每个单词的嵌入向量来形成虚拟句子,并用训练数据对其进行扩充。通过这种方式,该模型在保持原始干净数据性能的同时,对对手攻击具有较强的健壮性。DNE与网络架构无关,可扩展到用于NLP应用的大型模型(例如,BERT)。通过广泛的实验,我们证明了我们的方法在不同的网络体系结构和多个数据集上一致地比最近提出的防御方法有显著的优势。
Although deep neural networks have achieved prominent performance on many NLP tasks, they are vulnerable to adversarial examples. We propose Dirichlet Neighborhood Ensemble (DNE), a randomized method for training a robust model to defense synonym substitution-based attacks. During training, DNE forms virtual sentences by sampling embedding vectors for each word in an input sentence from a convex hull spanned by the word and its synonyms, and it augments them with the training data. In such a way, the model is robust to adversarial attacks while maintaining the performance on the original clean data. DNE is agnostic to the network architectures and scales to large models (e.g., BERT) for NLP applications. Through extensive experimentation, we demonstrate that our method consistently outperforms recently proposed defense methods by a significant margin across different network architectures and multiple data sets.