Fundamental limits on adversarial robustness

Fundamental limits on adversarial robustness
复制标题

对抗稳健性的基本限制

DOI:
--
复制
发表时间:
2015
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
P. Frossard
P. Frossard
中科院分区:
--
文献类型:
--
作者:
Alhussein Fawzi;Omar Fawzi;P. Frossard

文献摘要

被引文献

相似文献

本文的目的是分析最近在深度网络中发现的一个有趣现象,即它们对对抗性扰动的不稳定性(Szegedy et al., 2014)。我们提供了一个理论框架来分析分类器对对抗扰动的鲁棒性,并根据类之间的可区分性度量建立了一些分类器鲁棒性的基本限制。我们的结果表明,在涉及小可区分性的任务中,即使达到了良好的精度,所考虑的集中也没有分类器对对抗性扰动具有鲁棒性。此外,我们还证明了分类器对随机噪声的鲁棒性与其对对抗性扰动的鲁棒性之间存在明显的区别。具体来说,在高维情况下,对于线性分类器,前者比后者大得多。这一结果为两种鲁棒性之间的差异提供了理论解释,这是在神经网络背景下(Szegedy等人,2014)中经验观察到的。我们的理论框架表明,对抗性不稳定性是一种超越深度网络的现象,并影响所有分类器。与最初认为对抗示例是由神经网络的高度非线性引起的不同,我们的研究结果表明,与分类任务的难度相比,这种现象是由于分类器的低灵活性,这是由可分辨性度量捕获的。我们相信这些结果是ICML 2015年深度学习研讨会,里尔,法国。2015年版权归作者所有。有助于更好地理解对抗性不稳定性现象,以达到设计鲁棒分类器的目标。
The goal of this paper is to analyze an intriguing phenomenon recently discovered in deep networks, that is their instability to adversarial perturbations (Szegedy et al., 2014). We provide a theoretical framework for analyzing the robustness of classifiers to adversarial perturbations, and establish fundamental limits on the robustness of some classifiers in terms of a distinguishability measure between the classes. Our result implies that in tasks involving small distinguishability, no classifier in the considered set will be robust to adversarial perturbations, even if a good accuracy is achieved. Moreover, we show the existence of a clear distinction between the robustness of a classifier to random noise and its robustness to adversarial perturbations. Specifically, in high dimensions, the former is shown to be much larger than the latter for linear classifiers. This result gives a theoretical explanation for the discrepancy between the two robustness properties, which was empirically observed in (Szegedy et al., 2014) in the context of neural networks. Our theoretical framework shows that the adversarial instability is a phenomenon that goes beyond deep networks, and affects all classifiers. Unlike the initial belief that adversarial examples are caused by the high non-linearity of neural networks, our results suggest instead that this phenomenon is due to the low flexibility of classifiers, compared to the difficulty of the classification task, which is captured by the distinguishability measure. We believe these results ICML 2015 Workshop on Deep Learning, Lille, France. Copyright 2015 by the author(s). contribute to a better understanding of the phenomenon of adversarial instability to reach the goal of designing robust classifiers.