Communication-Efficient Stochastic Gradient Descent, with Applications to Neural Networks

Communication-Efficient Stochastic Gradient Descent, with Applications to Neural Networks
复制标题

通信高效的随机梯度下降及其在神经网络中的应用

DOI:
--
复制
发表时间:
2017
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
M. Vojnović
M. Vojnović
中科院分区:
--
文献类型:
--
作者:
Dan Alistarh;Demjan Grubic;Jerry Liu;Ryota Tomioka;M. Vojnović

文献摘要

被引文献

相似文献

随机梯度下降(SGD)的并行实现,由于其出色的可扩展性,已收到显着的研究关注。并行化SGD时的一个基本障碍是节点之间通信梯度更新的高带宽成本;因此,已经提出了几种有损的并行化算法,通过这些算法,节点只通信量化的梯度。虽然这些方法在实践中行之有效,但并不总是保证趋同,而且还不清楚是否可以改进。在本文中,我们提出了量化SGD(QSGD),一个家庭的梯度更新,提供收敛保证的压缩方案。QSGD允许用户平滑地权衡通信带宽和收敛时间:节点可以调整每次迭代发送的比特数,代价是可能更高的方差。我们表明,这种权衡是固有的,在这个意义上说,提高它过去的一些阈值将违反信息理论的下限。QSGD保证凸和非凸目标的收敛性,并可以扩展到随机方差减少技术。当应用于训练用于图像分类和自动语音识别的深度神经网络时,QSGD可以显著减少端到端训练时间。例如,在16个GPU上,我们可以在ImageNet上训练ResNet 152网络达到全精度,比全精度变体快1.8倍。
Parallel implementations of stochastic gradient descent (SGD) have received significant research attention, thanks to its excellent scalability properties. A fundamental barrier when parallelizing SGD is the high bandwidth cost of communicating gradient updates between nodes; consequently, several lossy compresion heuristics have been proposed, by which nodes only communicate quantized gradients. Although effective in practice, these heuristics do not always guarantee convergence, and it is not clear whether they can be improved. In this paper, we propose Quantized SGD (QSGD), a family of compression schemes for gradient updates which provides convergence guarantees. QSGD allows the user to smoothly trade off emph{communication bandwidth} and emph{convergence time}: nodes can adjust the number of bits sent per iteration, at the cost of possibly higher variance. We show that this trade-off is inherent, in the sense that improving it past some threshold would violate information-theoretic lower bounds. QSGD guarantees convergence for convex and non-convex objectives, under asynchrony, and can be extended to stochastic variance-reduced techniques. When applied to training deep neural networks for image classification and automated speech recognition, QSGD leads to significant reductions in end-to-end training time. For example, on 16GPUs, we can train the ResNet152 network to full accuracy on ImageNet 1.8x faster than the full-precision variant.