Efficient and Robust Classification for Sparse Attacks

Efficient and Robust Classification for Sparse Attacks
复制标题

DOI:
10.1109/isit50566.2022.9834832
复制
发表时间:
2022-01
期刊:
2022 IEEE International Symposium on Information Theory (ISIT)
影响因子:
--
通讯作者:
M. Beliaev;Payam Delgosha;Hamed Hassani;Ramtin Pedarsani
M. Beliaev;Payam Delgosha;Hamed Hassani;Ramtin Pedarsani
中科院分区:
其他
文献类型:
--
作者:
M. Beliaev;Payam Delgosha;Hamed Hassani;Ramtin Pedarsani

文献摘要

相似文献

在过去的二十年里,我们看到神经网络的普及与其分类准确性相结合。与此同时,我们也目睹了同样的预测模型是多么脆弱:对输入的微小扰动可能导致整个数据集的错误分类。在本文中,我们考虑的扰动有界的100-范数,这已被证明是有效的攻击领域的图像识别,自然语言处理和恶意软件检测。为此,我们提出了一种新的防御方法,包括“截断”和“对抗训练”。然后,我们从理论上研究高斯混合设置,并证明我们提出的分类器的渐近最优性。受我们获得的见解的启发,我们将这些组件扩展到神经网络分类器。我们使用MNIST和CIFAR数据集在计算机视觉领域进行数值实验,证明了神经网络的鲁棒分类错误的显着改善。
In the past two decades we have seen the popularity of neural networks increase in conjunction with their classification accuracy. Parallel to this, we have also witnessed how fragile the very same prediction models are: tiny perturbations to the inputs can cause misclassification errors throughout entire datasets. In this paper, we consider perturbations bounded by the ℓ0–norm, which have been shown as effective attacks in the domains of image-recognition, natural language processing, and malware-detection. To this end, we propose a novel defense method that consists of "truncation" and "adversarial training". We then theoretically study the Gaussian mixture setting and prove the asymptotic optimality of our proposed classifier. Motivated by the insights we obtain, we extend these components to neural network classifiers. We conduct numerical experiments in the domain of computer vision using the MNIST and CIFAR datasets, demonstrating significant improvement for the robust classification error of neural networks.