Demystifying Parallel and Distributed Deep Learning: An In-depth Concurrency Analysis

Demystifying Parallel and Distributed Deep Learning: An In-depth Concurrency Analysis
复制标题

揭开并行和分布式深度学习的神秘面纱:深度并发分析

DOI:
10.1145/3320060
复制
发表时间:
2019-09-01
影响因子:
16.6
通讯作者:
Hoefler, Torsten
Hoefler, Torsten
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ben-Nun, Tal;Hoefler, Torsten

文献摘要

被引文献

相似文献

深度神经网络(DNN)正在成为现代计算应用中的重要工具。加速他们的训练是一项重大挑战,技术范围从分布式算法到低级电路设计。在本次调查中,我们从理论角度描述了该问题,然后介绍了其并行化的方法。我们介绍了 DNN 架构的趋势以及由此产生的对并行化策略的影响。然后,我们回顾并建模 DNN 中不同类型的并发性:从单个算子,到网络推理和训练中的并行性,再到分布式深度学习。我们讨论异步随机优化、分布式系统架构、通信方案和神经架构搜索。基于这些方法,我们推断出深度学习中并行性的潜在方向。
Deep Neural Networks (DNNs) are becoming an important tool in modern computing applications. Accelerating their training is a major challenge and techniques range from distributed algorithms to low-level circuit design. In this survey, we describe the problem from a theoretical perspective, followed by approaches for its parallelization. We present trends in DNN architectures and the resulting implications on parallelization strategies. We then review and model the different types of concurrency in DNNs: from the single operator, through parallelism in network inference and training, to distributed deep learning. We discuss asynchronous stochastic optimization, distributed system architectures, communication schemes, and neural architecture search. Based on those approaches, we extrapolate potential directions for parallelism in deep learning.