Fast and scalable all-optical network architecture for distributed deep learning

Fast and scalable all-optical network architecture for distributed deep learning
复制标题

用于分布式深度学习的快速且可扩展的全光网络架构

DOI:
--
复制
发表时间:
2024
影响因子:
5
通讯作者:
G. Rouskas
G. Rouskas
中科院分区:
计算机科学1区
文献类型:
--
作者:
Wenzhe Li;Guojun Yuan;Zhan Wang;Guangming Tan;Peiheng Zhang;G. Rouskas

文献摘要

参考文献

相似文献

随着训练模型和数据集规模的不断增加,网络通信已成为分布式深度学习训练的主要瓶颈。为了应对这一挑战,我们提出了一种光学分布式深度学习(ODDL)架构。 ODDL 利用快速且可扩展的全光网络架构来加速分布式训练。该架构的主要特点之一是其基于流的传输调度和快速重新配置。这使得 ODDL 能够为每个流量动态分配专用光路,从而实现低网络延迟和高网络利用率。此外,ODDL 通过使用 LCoS-WSS 技术重新配置光开关,为训练任务提供物理隔离和定制的网络资源。 ODDL 拓扑还使用可调谐收发器来适应时变的流量模式。为了实现光路的精确和细粒度调度,我们提出了一种有效的分布式控制方案,该方案产生最小的延迟开销。我们对现实世界痕迹的评估展示了 ODDL 的卓越性能。当采用 1024 个节点和 100 Gbps 带宽实施时,与传统的胖树电气网络和光子 SiP-Ring 架构相比,ODDL 将 VGG19 训练速度分别提高了 1.6 美元和 1.7 美元。我们进一步构建了一个四节点测试台,我们的实验表明,ODDL 可以实现与理想电气交换网络相当的训练时间。
With the ever-increasing size of training models and datasets, network communication has emerged as a major bottleneck in distributed deep learning training. To address this challenge, we propose an optical distributed deep learning (ODDL) architecture. ODDL utilizes a fast yet scalable all-optical network architecture to accelerate distributed training. One of the key features of the architecture is its flow-based transmit scheduling with fast reconfiguration. This allows ODDL to allocate dedicated optical paths for each traffic stream dynamically, resulting in low network latency and high network utilization. Additionally, ODDL provides physically isolated and tailored network resources for training tasks by reconfiguring the optical switch using LCoS-WSS technology. The ODDL topology also uses tunable transceivers to adapt to time-varying traffic patterns. To achieve accurate and fine-grained scheduling of optical circuits, we propose an efficient distributed control scheme that incurs minimal delay overhead. Our evaluation on real-world traces showcases ODDL’s remarkable performance. When implemented with 1024 nodes and 100 Gbps bandwidth, ODDL accelerates VGG19 training by $1.6 \times$ and $1.7 \times$ compared to conventional fat-tree electrical networks and photonic SiP-Ring architectures, respectively. We further build a four-node testbed, and our experiments show that ODDL can achieve comparable training time compared to that of an ideal electrical switching network.
光交换数据中心网络中的亚纳秒时钟和数据恢复
DOI: 10.1109/ecoc.2018.8535333
发表时间: 2018
期刊: --
影响因子: --
作者:
Clark K
通讯作者: Clark K
DOI: --
发表时间: 2019-10
期刊: ArXiv
影响因子: --
作者:
Guanhua Wang;S. Venkataraman;Amar Phanishayee;J. Thelin;Nikhil R. Devanur;I. Stoica
通讯作者: Guanhua Wang;S. Venkataraman;Amar Phanishayee;J. Thelin;Nikhil R. Devanur;I. Stoica
DOI: --
发表时间: 2022-02
期刊: --
影响因子: --
作者:
Weiyang Wang;Moein Khazraee;Zhizhen Zhong;M. Ghobadi;Zhihao Jia;Dheevatsa Mudigere;Ying Zhang;
通讯作者: Weiyang Wang;Moein Khazraee;Zhizhen Zhong;M. Ghobadi;Zhihao Jia;Dheevatsa Mudigere;Ying Zhang;
使用时钟相位缓存的光交换数据中心的同步亚纳秒时钟和数据恢复
DOI: 10.1038/s41928-020-0423-y
发表时间: 2020
期刊: Nature Electronics
影响因子: 34.3
作者:
Clark K
通讯作者: Clark K
DOI: 10.1109/ccgrid51090.2021.00021
发表时间: 2021-05
期刊: 2021 IEEE/ACM 21st International Symposium on Cluster, Cloud and Internet Computing (CCGrid)
影响因子: --
作者:
Kawthar Shafie Khorassani;Ching-Hsiang Chu;Quentin G. Anthony;H. Subramoni;D. Panda
通讯作者: Kawthar Shafie Khorassani;Ching-Hsiang Chu;Quentin G. Anthony;H. Subramoni;D. Panda