Generating Natural Language Adversarial Examples

Generating Natural Language Adversarial Examples
复制标题

DOI:
10.18653/v1/d18-1316
复制
发表时间:
2018-04
期刊:
ArXiv
影响因子:
--
通讯作者:
M. Alzantot;Yash Sharma;Ahmed Elgohary;Bo-Jhang Ho;M. Srivastava;Kai-Wei Chang
M. Alzantot;Yash Sharma;Ahmed Elgohary;Bo-Jhang Ho;M. Srivastava;Kai-Wei Chang
中科院分区:
其他
文献类型:
--
作者:
M. Alzantot;Yash Sharma;Ahmed Elgohary;Bo-Jhang Ho;M. Srivastava;Kai-Wei Chang

文献摘要

被引文献

相似文献

深度神经网络(dnn)容易受到对抗性样本的影响,对正确分类的样本的扰动可能导致模型错误分类。在图像域,这些扰动通常对人类的感知几乎无法区分,导致人类和最先进的模型不一致。然而,在自然语言领域,微小的扰动是明显可感知的,并且替换单个单词可以彻底改变文档的语义。鉴于这些挑战,我们使用基于黑盒群体的优化算法来生成语义和语法相似的对抗性示例,这些示例分别以97%和70%的成功率骗过训练有素的情感分析和文本蕴含模型。我们还证明,92.3%的成功的情感分析对抗性示例被20个人类注释者分类到它们的原始标签上,并且这些示例明显非常相似。最后,我们讨论了使用对抗性训练作为防御的尝试,但未能产生改进,展示了我们对抗性示例的强度和多样性。我们希望我们的发现能鼓励研究人员在自然语言领域提高dnn的鲁棒性。
Deep neural networks (DNNs) are vulnerable to adversarial examples, perturbations to correctly classified examples which can cause the model to misclassify. In the image domain, these perturbations can often be made virtually indistinguishable to human perception, causing humans and state-of-the-art models to disagree. However, in the natural language domain, small perturbations are clearly perceptible, and the replacement of a single word can drastically alter the semantics of the document. Given these challenges, we use a black-box population-based optimization algorithm to generate semantically and syntactically similar adversarial examples that fool well-trained sentiment analysis and textual entailment models with success rates of 97% and 70%, respectively. We additionally demonstrate that 92.3% of the successful sentiment analysis adversarial examples are classified to their original label by 20 human annotators, and that the examples are perceptibly quite similar. Finally, we discuss an attempt to use adversarial training as a defense, but fail to yield improvement, demonstrating the strength and diversity of our adversarial examples. We hope our findings encourage researchers to pursue improving the robustness of DNNs in the natural language domain.