Graph interpolating activation improves both natural and robust accuracies in data-efficient deep learning

Graph interpolating activation improves both natural and robust accuracies in data-efficient deep learning
复制标题

DOI:
10.1017/s0956792520000406
复制
发表时间:
2019-07
影响因子:
1.9
通讯作者:
Bao Wang;S. Osher
Bao Wang;S. Osher
中科院分区:
数学4区
文献类型:
--
作者:
Bao Wang;S. Osher

文献摘要

相似文献

提高深度神经网络(dnn)的准确性和鲁棒性,并使其适应小数据训练是深度学习(DL)研究的主要任务。在本文中,我们用基于图拉普拉斯的高维插值函数代替dnn的输出激活函数,通常是与数据无关的softmax函数,该函数在连续极限下收敛于高维流形上的拉普拉斯-贝尔特拉米方程的解。此外,我们为这种新架构提出了端到端训练和测试算法。基于图插值激活的深度神经网络融合了深度学习和流形学习的优点。与使用softmax函数作为输出激活的传统dnn相比,新框架显示出以下主要优势:首先,它更适用于数据高效学习,在这种学习中,我们可以在不使用大量训练数据的情况下训练高容量dnn。其次,它显著提高了白盒攻击和黑盒攻击对干净图像的自然精度和对敌对图像的鲁棒精度。第三,它是半监督学习的自然选择。这篇论文是我们在2018年发表在NeurIPS上的早期工作的重要延伸。为了再现性,代码可在https://github.com/BaoWangMath/DNN-DataDependentActivation上获得。
Improving the accuracy and robustness of deep neural nets (DNNs) and adapting them to small training data are primary tasks in deep learning (DL) research. In this paper, we replace the output activation function of DNNs, typically the data-agnostic softmax function, with a graph Laplacian-based high-dimensional interpolating function which, in the continuum limit, converges to the solution of a Laplace–Beltrami equation on a high-dimensional manifold. Furthermore, we propose end-to-end training and testing algorithms for this new architecture. The proposed DNN with graph interpolating activation integrates the advantages of both deep learning and manifold learning. Compared to the conventional DNNs with the softmax function as output activation, the new framework demonstrates the following major advantages: First, it is better applicable to data-efficient learning in which we train high capacity DNNs without using a large number of training data. Second, it remarkably improves both natural accuracy on the clean images and robust accuracy on the adversarial images crafted by both white-box and black-box adversarial attacks. Third, it is a natural choice for semi-supervised learning. This paper is a significant extension of our earlier work published in NeurIPS, 2018. For reproducibility, the code is available at https://github.com/BaoWangMath/DNN-DataDependentActivation.