AdvMind: Inferring Adversary Intent of Black-Box Attacks

AdvMind: Inferring Adversary Intent of Black-Box Attacks
复制标题

DOI:
10.1145/3394486.3403241
复制
发表时间:
2020-06
期刊:
Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Ren Pang;Xinyang Zhang-;S. Ji;Xiapu Luo;Ting Wang
Ren Pang;Xinyang Zhang-;S. Ji;Xiapu Luo;Ting Wang
中科院分区:
其他
文献类型:
--
作者:
Ren Pang;Xinyang Zhang-;S. Ji;Xiapu Luo;Ting Wang

文献摘要

被引文献

相似文献

深度神经网络(DNN)本质上容易受到对抗性攻击,即使在黑盒设置下也是如此,在黑盒设置下,对手只能查询目标模型。在实践中,虽然可以有效地检测这种攻击(例如,观察大量相似但不相同的查询),准确地推断对手意图(例如,对手试图制造的敌对示例的目标类别),尤其是在攻击的早期阶段,这对于在许多情况下执行有效的威慑和补救威胁至关重要。在本文中,我们提出了AdvMind,这是一类新的估计模型,可以以鲁棒且迅速的方式推断黑盒对抗性攻击的对手意图。具体来说,为了实现鲁棒检测,AdvMind考虑了对手的适应性,使得她隐藏目标的尝试将显著增加攻击成本(例如,在查询的数量方面);为了实现及时的检测,AdvMind主动合成合理的查询结果,以从最大程度地暴露其意图的对手那里征求后续查询。通过对基准数据集和最先进的黑盒攻击进行广泛的实证评估,我们证明,AdvMind在观察不到3个查询批次后,平均以超过75%的准确率检测到对手意图,同时将自适应攻击的成本增加了60%以上。我们进一步讨论了AdvMind和其他防御方法对抗黑盒对抗攻击之间可能的协同作用,指出了几个有前途的研究方向。
Deep neural networks (DNNs) are inherently susceptible to adversarial attacks even under black-box settings, in which the adversary only has query access to the target models. In practice, while it may be possible to effectively detect such attacks (e.g., observing massive similar but non-identical queries), it is often challenging to exactly infer the adversary intent (e.g., the target class of the adversarial example the adversary attempts to craft) especially during early stages of the attacks, which is crucial for performing effective deterrence and remediation of the threats in many scenarios. In this paper, we present AdvMind, a new class of estimation models that infer the adversary intent of black-box adversarial attacks in a robust and prompt manner. Specifically, to achieve robust detection, AdvMind accounts for the adversary adaptiveness such that her attempt to conceal the target will significantly increase the attack cost (e.g., in terms of the number of queries); to achieve prompt detection, AdvMind proactively synthesizes plausible query results to solicit subsequent queries from the adversary that maximally expose her intent. Through extensive empirical evaluation on benchmark datasets and state-of-the-art black-box attacks, we demonstrate that on average AdvMind detects the adversary intent with over 75% accuracy after observing less than 3 query batches and meanwhile increases the cost of adaptive attacks by over 60%. We further discuss the possible synergy between AdvMind and other defense methods against black-box adversarial attacks, pointing to several promising research directions.