Accelerated Gradient-free Neural Network Training by Multi-convex Alternating Optimization
Accelerated Gradient-free Neural Network Training by Multi-convex Alternating Optimization
复制标题
DOI:
10.1016/j.neucom.2022.02.039
复制
发表时间:
2018-11
期刊:
影响因子:
6
通讯作者:
Junxiang Wang;Fuxun Yu;Xiangyi Chen;Liang Zhao
中科院分区:
文献类型:
--
作者:
Junxiang Wang;Fuxun Yu;Xiangyi Chen;Liang Zhao
In recent years, even though Stochastic Gradient Descent (SGD) and its variants are well-known for training neural networks, it suffers from limitations such as the lack of theoretical guarantees, vanishing gradients, and excessive sensitivity to input. To overcome these drawbacks, alternating minimization methods have attracted fast-increasing attention recently. As an emerging and open domain, however, several new challenges need to be addressed, including 1) Convergence properties are sensitive to penalty parameters, and 2) Slow theoretical convergence rate. We, therefore, propose a novel monotonous Deep Learning Alternating Minimization (mDLAM) algorithm to deal with these two challenges. Our innovative inequality-constrained formulation infinitely approximates the original problem with non-convex equality constraints, enabling our convergence proof of the proposed mDLAM algorithm regardless of the choice of hyperparameters. Our mDLAM algorithm is shown to achieve a fast linear convergence by the Nesterov acceleration technique. Extensive experiments on multiple benchmark datasets demonstrate the convergence, effectiveness, and efficiency of the proposed mDLAM algorithm.