SiP-ML: high-bandwidth optical network interconnects for machine learning training

SiP-ML: high-bandwidth optical network interconnects for machine learning training
复制标题

DOI:
10.1145/3452296.3472900
复制
发表时间:
2021-08
期刊:
Proceedings of the 2021 ACM SIGCOMM 2021 Conference
影响因子:
--
通讯作者:
Mehrdad Khani Shirkoohi;M. Ghobadi;M. Alizadeh;Ziyi Zhu;M. Glick;K. Bergman;A. Vahdat;Benjamin Klenk-Ben
Mehrdad Khani Shirkoohi;M. Ghobadi;M. Alizadeh;Ziyi Zhu;M. Glick;K. Bergman;A. Vahdat;Benjamin Klenk-Ben
中科院分区:
其他
文献类型:
--
作者:
Mehrdad Khani Shirkoohi;M. Ghobadi;M. Alizadeh;Ziyi Zhu;M. Glick;K. Bergman;A. Vahdat;Benjamin Klenk-Ben

文献摘要

被引文献

相似文献

本文提出光网络互连作为构建具有强大扩展特性的高带宽机器学习训练集群的关键推动者。我们的设计称为 SiP-ML,使用硅光子链路加速流行 DNN 模型的训练时间,该链路能够为每个 GPU 提供每秒多个太比特的带宽。 SiP-ML 通过混合数据和模型并行性在 GPU 之间划分训练作业,同时确保网络互连上可以有效支持通信模式。我们开发了任务划分和设备放置方法,将光互连的程度和重新配置延迟考虑在内。使用真实 DNN 模型进行的模拟表明,与最先进的电气网络相比,我们的方法将训练时间缩短了 1.3--9.1 倍。
This paper proposes optical network interconnects as a key enabler for building high-bandwidth ML training clusters with strong scaling properties. Our design, called SiP-ML, accelerates the training time of popular DNN models using silicon photonics links capable of providing multiple terabits-per-second of bandwidth per GPU. SiP-ML partitions the training job across GPUs with hybrid data and model parallelism while ensuring the communication pattern can be supported efficiently on the network interconnect. We develop task partitioning and device placement methods that take the degree and reconfiguration latency of optical interconnects into account. Simulations using real DNN models show that, compared to the state-of-the-art electrical networks, our approach improves training time by 1.3--9.1x.