Discrete Adversarial Attacks and Submodular Optimization with Applications to Text Classification

Discrete Adversarial Attacks and Submodular Optimization with Applications to Text Classification
复制标题

DOI:
--
复制
发表时间:
2018-12
期刊:
arXiv: Learning
影响因子:
--
通讯作者:
Qi Lei;Lingfei Wu;Pin-Yu Chen;A. Dimakis;I. Dhillon;M. Witbrock
Qi Lei;Lingfei Wu;Pin-Yu Chen;A. Dimakis;I. Dhillon;M. Witbrock
中科院分区:
其他
文献类型:
--
作者:
Qi Lei;Lingfei Wu;Pin-Yu Chen;A. Dimakis;I. Dhillon;M. Witbrock

文献摘要

被引文献

相似文献

对抗性示例是对输入的精心构造的修改,这些修改完全改变了分类器的输出,但人类无法察觉。尽管针对连续数据(例如图像和音频样本)的攻击取得了成功,但事实证明,为离散结构(例如文本)生成对抗性示例更具挑战性。在本文中,我们将针对集合函数的离散输入的攻击表述为优化任务。我们证明该集合函数对于一些流行的神经网络文本分类器在简化假设下是子模的。这一发现保证了使用贪婪算法的攻击具有 $1-1/e$ 近似因子。同时,我们展示了如何使用受攻击分类器的梯度来指导贪婪搜索。我们提出的优化方案的实证研究表明,在不同基线上的三种不同文本分类任务上,攻击能力和效率显着提高。我们还使用联合句子和单词释义技术来保持文本的原始语义和语法。这是通过人类受试者对我们生成的对抗性文本的质量和语义一致性的主观指标评估来验证的。
Adversarial examples are carefully constructed modifications to an input that completely change the output of a classifier but are imperceptible to humans. Despite these successful attacks for continuous data (such as image and audio samples), generating adversarial examples for discrete structures such as text has proven significantly more challenging. In this paper we formulate the attacks with discrete input on a set function as an optimization task. We prove that this set function is submodular for some popular neural network text classifiers under simplifying assumption. This finding guarantees a $1-1/e$ approximation factor for attacks that use the greedy algorithm. Meanwhile, we show how to use the gradient of the attacked classifier to guide the greedy search. Empirical studies with our proposed optimization scheme show significantly improved attack ability and efficiency, on three different text classification tasks over various baselines. We also use a joint sentence and word paraphrasing technique to maintain the original semantics and syntax of the text. This is validated by a human subject evaluation in subjective metrics on the quality and semantic coherence of our generated adversarial text.