Fast convergence rates of deep neural networks for classification

Fast convergence rates of deep neural networks for classification
复制标题

DOI:
10.1016/j.neunet.2021.02.012
复制
发表时间:
2021-03-04
期刊:
影响因子:
7.8
通讯作者:
Kim, Dongha
Kim, Dongha
中科院分区:
计算机科学1区
文献类型:
--
作者:
Kim, Yongdai;Ohn, Ilsang;Kim, Dongha

文献摘要

被引文献

相似文献

利用铰链损失学习的整流线性单元(ReLU)激活函数,推导出深度神经网络(DNN)分类器的快速收敛速率。我们考虑了真模型的三种情况:(1)光滑的决策边界,(2)光滑的条件类概率,(3)裕度条件(即,在决策边界附近的输入概率很小)。我们表明,使用铰链损失学习的DNN分类器在所有三种情况下都实现了快速收敛,前提是结构(即层数,节点数和稀疏性)是仔细选择的。一个重要的含义是,深度神经网络架构非常灵活,可以在各种情况下使用,而无需进行太多修改。此外,我们考虑了一种通过最小化交叉熵来学习的DNN分类器,并证明了在噪声指数和边际指数较大的条件下,DNN分类器具有较快的收敛速度。尽管它们很强大,但我们解释这两个条件对于图像分类问题来说并不是太荒谬。为了证实我们的理论解释,我们提出了一项小型数值研究的结果,以比较铰链损失和交叉熵。(c) 2021提交人。Elsevier Ltd.出版。
We derive the fast convergence rates of a deep neural network (DNN) classifier with the rectified linear unit (ReLU) activation function learned using the hinge loss. We consider three cases for a true model: (1) a smooth decision boundary, (2) smooth conditional class probability, and (3) the margin condition (i.e., the probability of inputs near the decision boundary is small). We show that the DNN classifier learned using the hinge loss achieves fast rate convergences for all three cases provided that the architecture (i.e., the number of layers, number of nodes and sparsity) is carefully selected. An important implication is that DNN architectures are very flexible for use in various cases without much modification. In addition, we consider a DNN classifier learned by minimizing the cross-entropy, and show that the DNN classifier achieves a fast convergence rate under the conditions that the noise exponent and margin exponent are large. Even though they are strong, we explain that these two conditions are not too absurd for image classification problems. To confirm our theoretical explanation, we present the results of a small numerical study conducted to compare the hinge loss and cross-entropy. (c) 2021 The Author(s). Published by Elsevier Ltd.