EventGraD: Event-Triggered Communication in Parallel Machine Learning

EventGraD: Event-Triggered Communication in Parallel Machine Learning
复制标题

DOI:
10.1016/j.neucom.2021.08.143
复制
发表时间:
2021-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Soumyadip Ghosh;B. Aquino;V. Gupta
Soumyadip Ghosh;B. Aquino;V. Gupta
中科院分区:
其他
文献类型:
--
作者:
Soumyadip Ghosh;B. Aquino;V. Gupta

文献摘要

相似文献

并行系统中的通信会带来巨大的开销,这通常会成为并行机器学习的瓶颈。为了减轻一些开销,在本文中,我们提出了 EventGraD - 一种具有事件触发通信的算法,用于并行机器学习中的随机梯度下降。该算法的主要思想是将并行机器学习中随机梯度下降的标准实现中每次迭代时的通信要求修改为仅在某些迭代中必要时才进行通信。我们提供了我们提出的算法的收敛性的理论分析。我们还实现了所提出的算法,对用于训练 CIFAR-10 数据集的流行残差神经网络进行数据并行训练,并表明 EventGraD 可以将通信负载减少多达 60%,同时保持相同水平的精度。此外,EventGraD 可以与 Top-K 稀疏化等其他方法结合使用,以进一步减少通信,同时保持准确性。
Communication in parallel systems imposes significant overhead which often turns out to be a bottleneck in parallel machine learning. To relieve some of this overhead, in this paper, we present EventGraD - an algorithm with event-triggered communication for stochastic gradient descent in parallel machine learning. The main idea of this algorithm is to modify the requirement of communication at every iteration in standard implementations of stochastic gradient descent in parallel machine learning to communicating only when necessary at certain iterations. We provide theoretical analysis of convergence of our proposed algorithm. We also implement the proposed algorithm for data-parallel training of a popular residual neural network used for training the CIFAR-10 dataset and show that EventGraD can reduce the communication load by up to 60% while retaining the same level of accuracy. In addition, EventGraD can be combined with other approaches such as Top-K sparsification to decrease communication further while maintaining accuracy.