SAFER: A Structure-free Approach for Certified Robustness to Adversarial Word Substitutions

SAFER: A Structure-free Approach for Certified Robustness to Adversarial Word Substitutions
复制标题

DOI:
10.18653/v1/2020.acl-main.317
复制
发表时间:
2020-05
期刊:
--
影响因子:
--
通讯作者:
Mao Ye;Chengyue Gong;Qiang Liu
Mao Ye;Chengyue Gong;Qiang Liu
中科院分区:
其他
文献类型:
--
作者:
Mao Ye;Chengyue Gong;Qiang Liu

文献摘要

被引文献

相似文献

最先进的NLP模型通常会被人类不知道的转换所欺骗,例如同义词替换。出于安全原因,开发具有认证鲁棒性的模型至关重要,该模型可以证明保证预测不会被任何可能的同义词替换所改变。在这项工作中,我们提出了一个新的随机平滑技术的基础上,它构建了一个随机集成应用随机字替换输入句子,并利用集成的统计特性,可证明的鲁棒性认证的鲁棒性的认证方法。我们的方法简单且无结构,因为它只需要模型输出的黑盒查询,因此可以应用于任何预训练模型(如BERT)和任何类型的模型(世界级或子字级)。我们的方法显着优于最近的最先进的方法,在IMDB和亚马逊文本分类任务的认证鲁棒性。据我们所知,我们是第一个在BERT等大型系统上实现认证鲁棒性的工作,具有实际意义的认证准确性。
State-of-the-art NLP models can often be fooled by human-unaware transformations such as synonymous word substitution. For security reasons, it is of critical importance to develop models with certified robustness that can provably guarantee that the prediction is can not be altered by any possible synonymous word substitution. In this work, we propose a certified robust method based on a new randomized smoothing technique, which constructs a stochastic ensemble by applying random word substitutions on the input sentences, and leverage the statistical properties of the ensemble to provably certify the robustness. Our method is simple and structure-free in that it only requires the black-box queries of the model outputs, and hence can be applied to any pre-trained models (such as BERT) and any types of models (world-level or subword-level). Our method significantly outperforms recent state-of-the-art methods for certified robustness on both IMDB and Amazon text classification tasks. To the best of our knowledge, we are the first work to achieve certified robustness on large systems such as BERT with practically meaningful certified accuracy.