Toward Model Parallelism for Deep Neural Network based on Gradient-free ADMM Framework

Toward Model Parallelism for Deep Neural Network based on Gradient-free ADMM Framework
复制标题

DOI:
10.1109/icdm50108.2020.00068
复制
发表时间:
2020-09
期刊:
2020 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
通讯作者:
Junxiang Wang;Zheng Chai;Yue Cheng;Liang Zhao
Junxiang Wang;Zheng Chai;Yue Cheng;Liang Zhao
中科院分区:
其他
文献类型:
--
作者:
Junxiang Wang;Zheng Chai;Yue Cheng;Liang Zhao

文献摘要

相似文献

交替方向乘法(ADMM)最近被提出作为深度学习问题的随机梯度下降(SGD)的潜在替代优化器。这是因为ADMM可以解决梯度消失和条件差的问题。此外,它在许多大规模深度学习应用中表现出良好的可扩展性。然而,由于变量之间的层依赖性,仍然缺乏用于深度神经网络的并行ADMM计算框架。在本文中,我们提出了一种新的并行深度学习ADMM框架(pdADMM)来实现层并行性:神经网络的每一层中的参数可以并行地独立更新。在较弱的条件下,理论上证明了所提出的pdADMM收敛于临界点. pdADMM的收敛速度被证明是$o(1/k)$,其中$k$是迭代次数。在六个基准数据集上进行的大量实验表明,我们提出的pdADMM可以为训练大规模深度神经网络带来超过10倍的加速,并且优于大多数比较方法。我们的代码可在https://github.com/xianggebenben/pdADMM上获得。
Alternating Direction Method of Multipliers (ADMM) has recently been proposed as a potential alternative optimizer to the Stochastic Gradient Descent(SGD) for deep learning problems. This is because ADMM can solve gradient vanishing and poor conditioning problems. Moreover, it has shown good scalability in many large-scale deep learning applications. However, there still lacks a parallel ADMM computational framework for deep neural networks because of layer dependency among variables. In this paper, we propose a novel parallel deep learning ADMM framework (pdADMM) to achieve layer parallelism: parameters in each layer of neural networks can be updated independently in parallel. The convergence of the proposed pdADMM to a critical point is theoretically proven under mild conditions. The convergence rate of the pdADMM is proven to be $o(1/k)$ where $k$ is the number of iterations. Extensive experiments on six benchmark datasets demonstrated that our proposed pdADMM can lead to more than 10 times speedup for training large-scale deep neural networks, and outperformed most of the comparison methods. Our code is available at: https://github.com/xianggebenben/pdADMM.