SPADE: A Spectral Method for Black-Box Adversarial Robustness Evaluation

SPADE: A Spectral Method for Black-Box Adversarial Robustness Evaluation
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
--
影响因子:
--
通讯作者:
Wuxinlin Cheng;Chenhui Deng;Zhiqiang Zhao;Yaohui Cai;Zhiru Zhang;Zhuo Feng
Wuxinlin Cheng;Chenhui Deng;Zhiqiang Zhao;Yaohui Cai;Zhiru Zhang;Zhuo Feng
中科院分区:
其他
文献类型:
--
作者:
Wuxinlin Cheng;Chenhui Deng;Zhiqiang Zhao;Yaohui Cai;Zhiru Zhang;Zhuo Feng

文献摘要

被引文献

相似文献

介绍了一种黑盒谱方法,用于评估给定机器学习(ML)模型的对抗鲁棒性。我们的方法,命名为SPADE,利用双射距离映射之间的输入/输出图构造近似对应的输入/输出数据的流形。利用广义Courant-Fischer定理,提出了一种用于评估给定模型对抗鲁棒性的SPADE评分,并证明了该评分是流形环境下最佳Lipschitz常数的上界.为了揭示最不健壮的数据样本极易受到对抗性攻击,我们开发了一种利用主导广义特征向量的谱图嵌入过程。这个嵌入步骤允许为每个数据样本分配一个鲁棒性分数,可以进一步利用该分数进行更有效的对抗训练。我们的实验表明,所提出的SPADE方法导致有希望的经验结果的神经网络模型,对抗训练的MNIST和CIFAR-10数据集。
A black-box spectral method is introduced for evaluating the adversarial robustness of a given machine learning (ML) model. Our approach, named SPADE, exploits bijective distance mapping between the input/output graphs constructed for approximating the manifolds corresponding to the input/output data. By leveraging the generalized Courant-Fischer theorem, we propose a SPADE score for evaluating the adversarial robustness of a given model, which is proved to be an upper bound of the best Lipschitz constant under the manifold setting. To reveal the most non-robust data samples highly vulnerable to adversarial attacks, we develop a spectral graph embedding procedure leveraging dominant generalized eigenvectors. This embedding step allows assigning each data sample a robustness score that can be further harnessed for more effective adversarial training. Our experiments show the proposed SPADE method leads to promising empirical results for neural network models that are adversarially trained with the MNIST and CIFAR-10 data sets.