Sparsity-based Defense Against Adversarial Attacks on Linear Classifiers

Sparsity-based Defense Against Adversarial Attacks on Linear Classifiers
复制标题

DOI:
10.1109/isit.2018.8437638
复制
发表时间:
2018-01
期刊:
2018 IEEE International Symposium on Information Theory (ISIT)
影响因子:
--
通讯作者:
Zhinus Marzi;S. Gopalakrishnan;Upamanyu Madhow;Ramtin Pedarsani
Zhinus Marzi;S. Gopalakrishnan;Upamanyu Madhow;Ramtin Pedarsani
中科院分区:
其他
文献类型:
--
作者:
Zhinus Marzi;S. Gopalakrishnan;Upamanyu Madhow;Ramtin Pedarsani

文献摘要

被引文献

相似文献

深度神经网络代表了机器学习在越来越多领域的最新发展,包括视觉、语音和自然语言处理。然而,最近的工作提出了关于这种架构的鲁棒性的重要问题,通过表明有可能通过微小的,几乎察觉不到的扰动引起分类错误。这种“对抗性攻击”或“对抗性示例”的脆弱性已被证实是由于深度网络的过度线性。在本文中,我们研究了这种现象的线性分类器的设置,并表明,它是可能的,利用自然数据中的稀疏性,以打击$\ell_{\infty}$ -有界对抗扰动。具体来说,我们证明了一个稀疏化前端通过合奏平均分析,MNIST手写数字数据库的实验结果的有效性。据我们所知,这是第一个表明稀疏性为防御对抗性攻击提供了理论上严格的框架的工作。
Deep neural networks represent the state of the art in machine learning in a growing number of fields, including vision, speech and natural language processing. However, recent work raises important questions about the robustness of such architectures, by showing that it is possible to induce classification errors through tiny, almost imperceptible, perturbations. Vulnerability to such “adversarial attacks”, or “adversarial examples”, has been conjectured to be due to the excessive linearity of deep networks. In this paper, we study this phenomenon in the setting of a linear classifier, and show that it is possible to exploit sparsity in natural data to combat $\ell_{\infty}$ -bounded adversarial perturbations. Specifically, we demonstrate the efficacy of a sparsifying front end via an ensemble averaged analysis, and experimental results for the MNIST handwritten digit database. To the best of our knowledge, this is the first work to show that sparsity provides a theoretically rigorous framework for defense against adversarial attacks.