In-network Aggregation for Shared Machine Learning Clusters

In-network Aggregation for Shared Machine Learning Clusters
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Nadeen Gebara;M. Ghobadi;Paolo Costa
Nadeen Gebara;M. Ghobadi;Paolo Costa
中科院分区:
环境科学与生态学3区
文献类型:
--
作者:
Nadeen Gebara;M. Ghobadi;Paolo Costa

文献摘要

相似文献

我们提出了P ANAMA,一种新的网络聚合框架,用于在服务于各种工作的共享集群上进行分布式机器学习(ML)训练。PANAMA包括两个关键组件:(i)定制的网络内硬件加速器,其可以支持线速率的迭代点梯度聚合而不损害准确性;以及(ii)轻量级负载平衡和拥塞控制协议,其利用ML数据并行作业的独特通信模式来实现跨不同作业的网络资源的公平共享,同时确保长时间的高吞吐量。运行作业,短作业和其他对延迟敏感的流量的低延迟。我们使用基于FPGA的原型与10 Gbps的收发器和大规模的模拟评估P ANAMA的可行性。我们的模拟结果表明,P ANAMA减少了大的工作的平均训练时间的一个因素的1.34。更重要的是,通过大幅降低大型数据并行作业对网络的负载,P ANAMA也为非聚合数据流提供了显著的贝内,特别是延迟敏感的短数据流,将其99%的瓦片完成时间缩短了4.5倍。
We present P ANAMA , a novel in-network aggregation framework for distributed machine learning (ML) training on shared clusters serving a variety of jobs. P ANAMA comprises two key components: ( i ) a custom in-network hardware accelerator that can support floating-point gradient aggregation at line rate without compromising accuracy; and ( ii ) a lightweight load-balancing and congestion control protocol that exploits the unique communication patterns of ML data-parallel jobs to enable fair sharing of network resources across different jobs while ensuring high throughput for long-running jobs and low latency for short jobs and other latency-sensitive traffic. We evaluate the feasibility of P ANAMA using an FPGA-based prototype with 10 Gbps transceivers and large-scale simulations. Our simulation results demonstrate that P ANAMA decreases the average training time of large jobs by up to a factor of 1.34. More importantly, by drastically decreasing the load placed on the network by large data-parallel jobs, P ANAMA provides significant benefits to non-aggregation flows too, especially latency-sensitive short flows, reducing their 99%-tile completion time by up to 4.5 × .