SQuARM-SGD: Communication-Efficient Momentum SGD for Decentralized Optimization

SQuARM-SGD: Communication-Efficient Momentum SGD for Decentralized Optimization
复制标题

DOI:
10.1109/jsait.2021.3103920
复制
发表时间:
2020-05
期刊:
IEEE Journal on Selected Areas in Information Theory
影响因子:
--
通讯作者:
Navjot Singh;Deepesh Data;Jemin George;S. Diggavi
Navjot Singh;Deepesh Data;Jemin George;S. Diggavi
中科院分区:
其他
文献类型:
--
作者:
Navjot Singh;Deepesh Data;Jemin George;S. Diggavi

文献摘要

相似文献

在本文中,我们提出并分析了SQuARM-SGD,这是一种用于在网络上分散训练大规模机器学习模型的通信高效算法。在SQuARM-SGD中,每个节点使用Nesterov的动量执行固定数量的本地SGD步骤,然后将稀疏化和量化的更新发送到由本地可计算触发标准调节的邻居。我们为一般(非凸)和凸光滑目标提供了算法的收敛保证,据我们所知,这是第一次对压缩分散SGD进行动量更新的理论分析。我们证明了SQuARM-SGD的收敛速度与vanilla SGD相匹配。我们的经验表明,包括动量更新SQuARM-SGD可以导致更好的测试性能比目前的最先进的不考虑动量更新。
In this paper, we propose and analyze SQuARM-SGD, a communication-efficient algorithm for decentralized training of large-scale machine learning models over a network. In SQuARM-SGD, each node performs a fixed number of local SGD steps using Nesterov’s momentum and then sends sparsified and quantized updates to its neighbors regulated by a locally computable triggering criterion. We provide convergence guarantees of our algorithm for general (non-convex) and convex smooth objectives, which, to the best of our knowledge, is the first theoretical analysis for compressed decentralized SGD with momentum updates. We show that the convergence rate of SQuARM-SGD matches that of vanilla SGD. We empirically show that including momentum updates in SQuARM-SGD can lead to better test performance than the current state-of-the-art which does not consider momentum updates.