On Optimizing the Communication of Model Parallelism

On Optimizing the Communication of Model Parallelism
复制标题

DOI:
10.48550/arxiv.2211.05322
复制
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Yonghao Zhuang;Hexu Zhao;Lianmin Zheng;Zhuohan Li;Eric P. Xing;Qirong Ho;Joseph E. Gonzalez
Yonghao Zhuang;Hexu Zhao;Lianmin Zheng;Zhuohan Li;Eric P. Xing;Qirong Ho;Joseph E. Gonzalez
中科院分区:
其他
文献类型:
--
作者:
Yonghao Zhuang;Hexu Zhao;Lianmin Zheng;Zhuohan Li;Eric P. Xing;Qirong Ho;Joseph E. Gonzalez

文献摘要

相似文献

我们研究了大规模模型并行深度学习(DL)中一种新颖而重要的通信模式,我们称之为跨网格重分片。当模型并行的两个范例--操作符内并行和操作符间并行--结合起来支持大型集群上的大型模型时,就会出现这种模式。在跨网格重新分片中,分片的张量需要从源设备网格发送到目的地设备网格,在目的地设备网格上,张量可以以相同或不同的布局分布。我们将其形式化为多对多组播通信问题,并表明现有方法要么是次优的,要么不能推广到不同的网络拓扑结构或张量布局,这是由于不同的模型架构和并行策略。然后,我们提出了两个贡献,以解决跨网格重新分片:一个有效的基于广播的通信系统,和一个“广播友好”的管道时间表。在微基准测试中,我们的整个系统在各种张量和网格布局上的性能比现有系统高出10倍。在GPT-3和U-Transformer两个大型模型的端到端训练中,我们分别将吞吐量提高了10%和50%。
We study a novel and important communication pattern in large-scale model-parallel deep learning (DL), which we call cross-mesh resharding. This pattern emerges when the two paradigms of model parallelism - intra-operator and inter-operator parallelism - are combined to support large models on large clusters. In cross-mesh resharding, a sharded tensor needs to be sent from a source device mesh to a destination device mesh, on which the tensor may be distributed with the same or different layouts. We formalize this as a many-to-many multicast communication problem, and show that existing approaches either are sub-optimal or do not generalize to different network topologies or tensor layouts, which result from different model architectures and parallelism strategies. We then propose two contributions to address cross-mesh resharding: an efficient broadcast-based communication system, and an"overlapping-friendly"pipeline schedule. On microbenchmarks, our overall system outperforms existing ones by up to 10x across various tensor and mesh layouts. On end-to-end training of two large models, GPT-3 and U-Transformer, we improve throughput by 10% and 50%, respectively.