Accordion: Adaptive Gradient Communication via Critical Learning Regime Identification

Accordion: Adaptive Gradient Communication via Critical Learning Regime Identification
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Saurabh Agarwal;Hongyi Wang;Kangwook Lee;S. Venkataraman;Dimitris Papailiopoulos
Saurabh Agarwal;Hongyi Wang;Kangwook Lee;S. Venkataraman;Dimitris Papailiopoulos
中科院分区:
其他
文献类型:
--
作者:
Saurabh Agarwal;Hongyi Wang;Kangwook Lee;S. Venkataraman;Dimitris Papailiopoulos

文献摘要

相似文献

由于计算节点之间频繁传输模型更新,分布式模型训练面临着通信瓶颈。为了缓解这些瓶颈,从业者使用稀疏化、量化或低秩更新等梯度压缩技术。这些技术通常需要选择静态压缩比,通常需要用户在模型精度和每次迭代加速之间进行权衡。在这项工作中,我们表明,由于选择高压缩比而导致的性能下降并不是根本性的。自适应压缩策略可以减少通信,同时保持最终测试的准确性。受到最近关于关键学习机制的发现的启发,其中小的梯度误差可能对模型性能产生不可恢复的影响,我们提出了 Accordion 一种简单而有效的自适应压缩算法。虽然 Accordion 平均保持足够高的压缩率,但它可以避免在关键学习状态下过度压缩梯度,这是通过简单的基于梯度范数的标准检测到的。我们对分布式环境中的大量机器学习任务进行的广泛实验研究表明,Accordion 保持了与未压缩训练类似的模型精度,但与静态方法相比,压缩效果提高了 5.5 倍,端到端加速提高了 4.1 倍。我们表明,Accordion 还可以调整批量大小,这是缓解通信瓶颈的另一种流行策略。
Distributed model training suffers from communication bottlenecks due to frequent model updates transmitted across compute nodes. To alleviate these bottlenecks, practitioners use gradient compression techniques like sparsification, quantization, or low-rank updates. The techniques usually require choosing a static compression ratio, often requiring users to balance the trade-off between model accuracy and per-iteration speedup. In this work, we show that such performance degradation due to choosing a high compression ratio is not fundamental. An adaptive compression strategy can reduce communication while maintaining final test accuracy. Inspired by recent findings on critical learning regimes, in which small gradient errors can have irrecoverable impact on model performance, we propose Accordion a simple yet effective adaptive compression algorithm. While Accordion maintains a high enough compression rate on average, it avoids over-compressing gradients whenever in critical learning regimes, detected by a simple gradient-norm based criterion. Our extensive experimental study over a number of machine learning tasks in distributed environments indicates that Accordion, maintains similar model accuracy to uncompressed training, yet achieves up to 5.5x better compression and up to 4.1x end-to-end speedup over static approaches. We show that Accordion also works for adjusting the batch size, another popular strategy for alleviating communication bottlenecks.