US-Byte: An Efficient Communication Framework for Scheduling Unequal-Sized Tensor Blocks in Distributed Deep Learning

US-Byte: An Efficient Communication Framework for Scheduling Unequal-Sized Tensor Blocks in Distributed Deep Learning
复制标题

DOI:
10.1109/tpds.2023.3331372
复制
发表时间:
2024-01
影响因子:
5.3
通讯作者:
Yunqi Gao;Bing Hu;Mahdi Boloursaz Mashhadi;A-Long Jin;Pei Xiao;Chunming Wu
Yunqi Gao;Bing Hu;Mahdi Boloursaz Mashhadi;A-Long Jin;Pei Xiao;Chunming Wu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yunqi Gao;Bing Hu;Mahdi Boloursaz Mashhadi;A-Long Jin;Pei Xiao;Chunming Wu

文献摘要

相似文献

通信瓶颈严重限制了分布式深度学习的可扩展性,有效的通信调度通过重叠计算和通信任务来加速分布式DNN训练。然而,现有的基于张量划分的方法效率不高,并且面临两个挑战:1)并行传输的张量块的固定数量不一定能最小化通信开销; 2)尽管优先传输靠近输入层的张量块的调度顺序可以在下一次迭代中更早地开始前向传播,但没有获得最短的每次迭代时间。在本文中,我们提出了一个高效的通信框架,称为US-Byte。它可以以接近最优的顺序调度大小不等的张量块,以最小化训练时间。我们分两个阶段建立了US-Byte的数学模型:1)梯度通信和反向传播的重叠,2)梯度通信和前向传播的重叠。我们从理论上推导出第二阶段的最佳解决方案,并有效地解决了第一阶段的低复杂度算法。我们在PyTorch框架上实现了US-Byte架构。在两个不同的8节点GPU集群上进行的大量实验表明,与Bytecom和WFBP相比,US-Byte可以分别实现高达1.26倍和1.56倍的加速比。我们进一步利用128个GPU的模拟来验证US-Byte的潜在扩展性能。模拟结果表明,与最先进的通信框架相比,US-Byte可以实现高达1.69倍的加速比。
The communication bottleneck severely constrains the scalability of distributed deep learning, and efficient communication scheduling accelerates distributed DNN training by overlapping computation and communication tasks. However, existing approaches based on tensor partitioning are not efficient and suffer from two challenges: 1) the fixed number of tensor blocks transferred in parallel can not necessarily minimize the communication overheads; 2) although the scheduling order that preferentially transmits tensor blocks close to the input layer can start forward propagation in the next iteration earlier, the shortest per-iteration time is not obtained. In this paper, we propose an efficient communication framework called US-Byte. It can schedule unequal-sized tensor blocks in a near-optimal order to minimize the training time. We build the mathematical model of US-Byte by two phases: 1) the overlap of gradient communication and backward propagation, and 2) the overlap of gradient communication and forward propagation. We theoretically derive the optimal solution for the second phase and efficiently solve the first phase with a low-complexity algorithm. We implement the US-Byte architecture on PyTorch framework. Extensive experiments on two different 8-node GPU clusters demonstrate that US-Byte can achieve up to 1.26x and 1.56x speedup compared to ByteScheduler and WFBP, respectively. We further exploit simulations of 128 GPUs to verify the potential scaling performance of US-Byte. Simulation results show that US-Byte can achieve up to 1.69x speedup compared to the state-of-the-art communication framework.