Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks

Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
复制标题

DOI:
10.1109/sp54263.2024.00053
复制
发表时间:
2023-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Xinyu Zhang;Hanbin Hong;Yuan Hong;Peng Huang;Binghui Wang;Zhongjie Ba;Kui Ren
Xinyu Zhang;Hanbin Hong;Yuan Hong;Peng Huang;Binghui Wang;Zhongjie Ba;Kui Ren
中科院分区:
其他
文献类型:
--
作者:
Xinyu Zhang;Hanbin Hong;Yuan Hong;Peng Huang;Binghui Wang;Zhongjie Ba;Kui Ren

文献摘要

相似文献

语言模型,特别是基本文本分类模型,已经被证明容易受到文本对抗性攻击,如同义词替换和单词插入攻击。为了防御这种攻击,越来越多的研究致力于提高模型的稳健性。然而,提供可证明的稳健性保证而不是经验稳健性仍然是广泛未被探索的。本文提出了一种基于随机化平滑的自然语言处理(NLP)广义认证健壮性框架Text-CRS。据我们所知,现有的NLP认证方案只能证明对同义词替换攻击中的$\ell_0$扰动的健壮性。将词级的敌意操作(即同义词替换、单词重排、插入和删除)表示为置换和嵌入变换的组合,提出了新颖的平滑定理,从而在置换空间和嵌入空间中推导出对此类敌意操作的稳健界。为了进一步提高认证精度和半径,我们考虑了离散字之间的数值关系,并选择合适的噪声分布进行随机化平滑。最后,我们在多种语言模型和数据集上进行了大量的实验。Text-CRS可以处理所有四种不同的词级对抗性操作,并实现显著的准确率提高。除了在同义词替换攻击方面优于最先进的认证外,我们还提供了首个关于四个词级操作的认证准确度和半径的基准。
The language models, especially the basic text classification models, have been shown to be susceptible to textual adversarial attacks such as synonym substitution and word insertion attacks. To defend against such attacks, a growing body of research has been devoted to improving the model robustness. However, providing provable robustness guarantees instead of empirical robustness is still widely unexplored. In this paper, we propose Text-CRS, a generalized certified robustness framework for natural language processing (NLP) based on randomized smoothing. To our best knowledge, existing certified schemes for NLP can only certify the robustness against $\ell_0$ perturbations in synonym substitution attacks. Representing each word-level adversarial operation (i.e., synonym substitution, word reordering, insertion, and deletion) as a combination of permutation and embedding transformation, we propose novel smoothing theorems to derive robustness bounds in both permutation and embedding space against such adversarial operations. To further improve certified accuracy and radius, we consider the numerical relationships between discrete words and select proper noise distributions for the randomized smoothing. Finally, we conduct substantial experiments on multiple language models and datasets. Text-CRS can address all four different word-level adversarial operations and achieve a significant accuracy improvement. We also provide the first benchmark on certified accuracy and radius of four word-level operations, besides outperforming the state-of-the-art certification against synonym substitution attacks.