Accelerated Gradient-free Neural Network Training by Multi-convex Alternating Optimization

Accelerated Gradient-free Neural Network Training by Multi-convex Alternating Optimization
复制标题

DOI:
10.1016/j.neucom.2022.02.039
复制
发表时间:
2018-11
期刊:
影响因子:
6
通讯作者:
Junxiang Wang;Fuxun Yu;Xiangyi Chen;Liang Zhao
Junxiang Wang;Fuxun Yu;Xiangyi Chen;Liang Zhao
中科院分区:
计算机科学2区
文献类型:
--
作者:
Junxiang Wang;Fuxun Yu;Xiangyi Chen;Liang Zhao

文献摘要

相似文献

近年来,尽管随机梯度下降(SGD)及其变体在训练神经网络方面非常有名,但它仍然存在诸如缺乏理论保证、梯度消失以及对输入过于敏感等局限性。为了克服这些缺点,交替最小化方法近年来引起了越来越多的关注。然而,作为一个新兴的开放领域,需要解决一些新的挑战,包括1)收敛性质对惩罚参数敏感;2)理论收敛速度慢。因此,我们提出了一种新颖的单调深度学习交替最小化(mDLAM)算法来处理这两个挑战。我们创新的不等式约束公式无限逼近原始问题的非凸等式约束,使我们提出的mDLAM算法的收敛性证明与超参数的选择无关。我们的mDLAM算法通过Nesterov加速技术实现了快速的线性收敛。在多个基准数据集上的大量实验证明了所提出的mDLAM算法的收敛性、有效性和高效性。
In recent years, even though Stochastic Gradient Descent (SGD) and its variants are well-known for training neural networks, it suffers from limitations such as the lack of theoretical guarantees, vanishing gradients, and excessive sensitivity to input. To overcome these drawbacks, alternating minimization methods have attracted fast-increasing attention recently. As an emerging and open domain, however, several new challenges need to be addressed, including 1) Convergence properties are sensitive to penalty parameters, and 2) Slow theoretical convergence rate. We, therefore, propose a novel monotonous Deep Learning Alternating Minimization (mDLAM) algorithm to deal with these two challenges. Our innovative inequality-constrained formulation infinitely approximates the original problem with non-convex equality constraints, enabling our convergence proof of the proposed mDLAM algorithm regardless of the choice of hyperparameters. Our mDLAM algorithm is shown to achieve a fast linear convergence by the Nesterov acceleration technique. Extensive experiments on multiple benchmark datasets demonstrate the convergence, effectiveness, and efficiency of the proposed mDLAM algorithm.