DLion: Decentralized Distributed Deep Learning in Micro-Clouds

DLion: Decentralized Distributed Deep Learning in Micro-Clouds
复制标题

DOI:
10.1145/3431379.3460643
复制
发表时间:
2020-06
期刊:
Proceedings of the 30th International Symposium on High-Performance Parallel and Distributed Computing
影响因子:
--
通讯作者:
Rankyung Hong;A. Chandra
Rankyung Hong;A. Chandra
中科院分区:
其他
文献类型:
--
作者:
Rankyung Hong;A. Chandra

文献摘要

相似文献

深度学习(DL)是一种流行的技术,用于从世界各地的边缘设备快速生成的大量数据(如图片,视频,消息)中构建模型。由于隐私、成本和性能原因,通过WAN将大量数据从边缘迁移到集中式数据中心进行培训通常是不可行的。同时,由于边缘设备的资源有限,在边缘设备上训练大型DL模型是不可行的。DL训练分布式数据的一个有吸引力的替代方案是使用微云--部署在多个位置的边缘设备附近的小规模云。然而,微云提出了计算和网络资源异构性以及动态性的挑战。在本文中,我们介绍了DLion,这是一种新的通用分散式分布式DL系统,旨在解决微云环境中的关键挑战,以减少整体训练时间并提高模型准确性。我们提出了DLion中的三个关键技术:(1)加权动态并行化,以最大限度地提高数据并行性,以处理异构和动态计算能力,(2)每链路优先梯度交换,以减少通信开销的基础上可用的网络容量模型更新,和(3)直接知识转移,以提高模型的准确性,通过合并最佳性能的模型参数。我们在TensorFlow之上构建了DLion的原型,并展示了DLion在Amazon GPU集群中实现了高达4.2倍的加速,在CPU集群中实现了高达2倍的加速,并在四个最先进的分布式DL系统中提高了26%的模型准确性。
Deep learning (DL) is a popular technique for building models from large quantities of data such as pictures, videos, messages generated from edges devices at rapid pace all over the world. It is often infeasible to migrate large quantities of data from the edges to centralized data center(s) over WANs for training due to privacy, cost, and performance reasons. At the same time, training large DL models on edge devices is infeasible due to their limited resources. An attractive alternative for DL training distributed data is to use micro-clouds---small-scale clouds deployed near edge devices in multiple locations. However, micro-clouds present the challenges of both computation and network resource heterogeneity as well as dynamism. In this paper, we introduce DLion, a new and generic decentralized distributed DL system designed to address the key challenges in micro-cloud environments, in order to reduce overall training time and improve model accuracy. We present three key techniques in DLion: (1) Weighted dynamic batching to maximize data parallelism for dealing with heterogeneous and dynamic compute capacity, (2) Per-link prioritized gradient exchange to reduce communication overhead for model updates based on available network capacity, and (3) Direct knowledge transfer to improve model accuracy by merging the best performing model parameters. We build a prototype of DLion on top of TensorFlow and show that DLion achieves up to 4.2X speedup in an Amazon GPU cluster, and up to 2X speed up and 26% higher model accuracy in a CPU cluster over four state-of-the-art distributed DL systems.