Blink: Fast and Generic Collectives for Distributed ML

Blink: Fast and Generic Collectives for Distributed ML
复制标题

DOI:
--
复制
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Guanhua Wang;S. Venkataraman;Amar Phanishayee;J. Thelin;Nikhil R. Devanur;I. Stoica
Guanhua Wang;S. Venkataraman;Amar Phanishayee;J. Thelin;Nikhil R. Devanur;I. Stoica
中科院分区:
其他
文献类型:
--
作者:
Guanhua Wang;S. Venkataraman;Amar Phanishayee;J. Thelin;Nikhil R. Devanur;I. Stoica

文献摘要

被引文献

相似文献

跨gpu的模型参数同步为大规模数据并行训练带来了很高的开销。面对日益增长的硬件异构性,现有的参数同步协议不能有效地利用可用的网络资源。为了解决这个问题,我们提出了Blink,一个通过打包生成树动态生成最优通信原语的集合通信库。我们提出了最小化生成树数量的技术,并扩展Blink以利用异构通信通道实现更快的数据传输。评估表明,与最先进的(NCCL)相比,Blink可以实现高达8倍的模型同步速度,并将图像分类任务的端到端训练时间减少多达40%。
Model parameter synchronization across GPUs introduces high overheads for data-parallel training at scale. Existing parameter synchronization protocols cannot effectively leverage available network resources in the face of ever increasing hardware heterogeneity. To address this, we propose Blink, a collective communication library that dynamically generates optimal communication primitives by packing spanning trees. We propose techniques to minimize the number of trees generated and extend Blink to leverage heterogeneous communication channels for faster data transfers. Evaluations show that compared to the state-of-the-art (NCCL), Blink can achieve up to 8× faster model synchronization, and reduce end-to-end training time for image classification tasks by up to 40%.