TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training Jobs

TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training Jobs
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
--
影响因子:
--
通讯作者:
Weiyang Wang;Moein Khazraee;Zhizhen Zhong;M. Ghobadi;Zhihao Jia;Dheevatsa Mudigere;Ying Zhang;
Weiyang Wang;Moein Khazraee;Zhizhen Zhong;M. Ghobadi;Zhihao Jia;Dheevatsa Mudigere;Ying Zhang;
中科院分区:
其他
文献类型:
--
作者:
Weiyang Wang;Moein Khazraee;Zhizhen Zhong;M. Ghobadi;Zhihao Jia;Dheevatsa Mudigere;Ying Zhang;

文献摘要

被引文献

相似文献

我们提出了TopoOpt,这是一种用于深度神经网络(DNN)训练工作负载的新型直连结构。TopoOpt在三个维度上协同优化分布式训练过程:计算、通信和网络拓扑。我们展示了AllReduce流量的可变性,并利用此属性为DNN训练作业构建高效的网络拓扑。TopoOpt然后使用交替优化技术和一种名为TotientPerms的群论启发算法来找到最佳网络拓扑和路由计划,以及并行化策略。我们构建了一个功能齐全的12节点直连原型,具有100 Gbps的远程直接内存访问(RDMA)转发。对真实的分布式训练模型的大规模模拟表明,与类似成本的Fat-Tree互连相比,TopoOpt将DNN训练时间减少了3.4倍。
We propose TopoOpt, a novel direct-connect fabric for deep neural network (DNN) training workloads. TopoOpt co-optimizes the distributed training process across three dimensions: computation, communication, and network topology. We demonstrate the mutability of AllReduce traffic, and leverage this property to construct efficient network topologies for DNN training jobs. TopoOpt then uses an alternating optimization technique and a group theory-inspired algorithm called TotientPerms to find the best network topology and routing plan, together with a parallelization strategy. We build a fully functional 12-node direct-connect prototype with remote direct memory access (RDMA) forwarding at 100 Gbps. Large-scale simulations on real distributed training models show that, compared to similar-cost Fat-Tree interconnects, TopoOpt reduces DNN training time by up to 3.4x.