Impact of Synchronization Topology on DML Performance: Both Logical Topology and Physical Topology

Impact of Synchronization Topology on DML Performance: Both Logical Topology and Physical Topology
复制标题

同步拓扑对 DML 性能的影响:逻辑拓扑和物理拓扑

DOI:
10.1109/tnet.2021.3117042
复制
发表时间:
2022-04
期刊:
IEEE/ACM TRANSACTIONS ON NETWORKING
影响因子:
--
通讯作者:
Dan Li
Dan Li
中科院分区:
其他
文献类型:
--
作者:
Shuai Wang;Jinkun Geng;Dan Li

文献摘要

相似文献

为了处理日益庞大的训练数据和模型,研究人员和工程师求助于数据中心中的多个服务器来进行分布式机器学习(DML)。一方面,DML使我们能够利用多个服务器的计算能力,从而有效地加速那些计算密集型任务。另一方面,由于这些服务器之间的参数同步,DML也产生了巨大的通信成本。在本文中,我们想要探讨同步拓扑,包括逻辑拓扑和物理拓扑,对DML性能的影响。首先,我们重新研究了现有的参数同步逻辑拓扑,如参数服务器和环ALL约简,我们发现这些平面同步拓扑在运行大规模DML训练时效率不高。因此,我们提出了一种分层的参数同步拓扑,称为HIPS,它即使在大范围内也能实现高效的参数同步。然后,比较了两种典型的物理网络拓扑,即Fat-Tree和BCube。根据我们的分析,BCube比Fat-Tree具有许多优势,如更高的带宽、更好的负载均衡和更低的硬件成本。仿真结果还表明,BCube对RDMA更加友好。依托HIPS和BCube的优势,“HIPS+BCube”的GST比其他组合低12%~70%。而且,当集群规模从16个增加到1024个时,HIPS+BCube的性能只下降了6.5%,而Ring+BCube的性能下降了44.6%。因此,我们认为“HIPS+BCube”是大规模造福DML的最佳解决方案。
To tackle the increasingly larger training data and models, researchers and engineers resort to multiple servers in a data center for distributed machine learning (DML). On one hand, DML enables us to leverage the computation power of multiple servers, which can effectively accelerate those computation-intensive tasks. On the other hand, DML also incurs significant communication cost due to parameter synchronization among these servers. In this paper, we want to explore the impact of synchronization topology, including both logical topology and physical topology, on the DML performance. First, we revisit the existing logical topologies, e.g., parameter server and ring allreduce, for parameter synchronization, and we find that these flat synchronization topologies is inefficient when running a large-scale DML training. Therefore, we propose a hierarchical parameter synchronization topology, called HiPS, which can achieve efficient parameter synchronization even on a large scale. Then, we compare two representative physical network topologies, namely, Fat-Tree and BCube. Based on our analyses, BCube has many advantages over Fat-Tree, e.g., higher bandwidth, better load balance, and lower hardware cost. The simulation results also show that BCube is more friendly to RDMA. Relying on the advantages of HiPS and BCube, the GST of “HiPS+BCube” is 12% ~ 70% lower than other combinations. Moreover, when the cluster size increases from 16 to 1024, the performance of “HiPS+BCube” only drops by 6.5%, while the performance of “Ring+BCube” drops by 44.6%. Hence, we believe “HiPS+BCube” is the optimal solution to benefit DML in large scale.