ATOMO: Communication-efficient Learning via Atomic Sparsification

ATOMO: Communication-efficient Learning via Atomic Sparsification
复制标题

DOI:
--
复制
发表时间:
2018-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Hongyi Wang;Scott Sievert;Shengchao Liu;Zachary B. Charles;Dimitris Papailiopoulos;S. Wright
Hongyi Wang;Scott Sievert;Shengchao Liu;Zachary B. Charles;Dimitris Papailiopoulos;S. Wright
中科院分区:
其他
文献类型:
--
作者:
Hongyi Wang;Scott Sievert;Shengchao Liu;Zachary B. Charles;Dimitris Papailiopoulos;S. Wright

文献摘要

被引文献

相似文献

由于计算节点之间频繁传输梯度更新,分布式模型训练会遭受通信开销。为了减轻这些开销,一些研究建议使用稀疏随机梯度。我们认为这些是通用稀疏化方法的各个方面,可以对任何可能的原子分解进行操作。值得注意的例子包括逐元素分解、奇异值分解和傅立叶分解。我们提出了 ATOMO,一个随机梯度原子稀疏​​化的通用框架。给定梯度、原子分解和稀疏预算,ATOMO 给出原子的随机无偏稀疏化,从而最小化方差。我们证明了 QSGD 和 TernGrad 等方法是 ATOMO 的特例,并表明在奇异值分解(SVD)中稀疏梯度(而不是坐标梯度)可以显着加快分布式训练的速度。
Distributed model training suffers from communication overheads due to frequent gradient updates transmitted between compute nodes. To mitigate these overheads, several studies propose the use of sparsified stochastic gradients. We argue that these are facets of a general sparsification method that can operate on any possible atomic decomposition. Notable examples include element-wise, singular value, and Fourier decompositions. We present ATOMO, a general framework for atomic sparsification of stochastic gradients. Given a gradient, an atomic decomposition, and a sparsity budget, ATOMO gives a random unbiased sparsification of the atoms minimizing variance. We show that methods such as QSGD and TernGrad are special cases of ATOMO and show that sparsifiying gradients in their singular value decomposition (SVD), rather than the coordinate-wise one, can lead to significantly faster distributed training.