HyPar: Towards Hybrid Parallelism for Deep Learning Accelerator Array

HyPar: Towards Hybrid Parallelism for Deep Learning Accelerator Array
复制标题

DOI:
10.1109/hpca.2019.00027
复制
发表时间:
2019-01
期刊:
2019 IEEE International Symposium on High Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Linghao Song;Jiachen Mao;Youwei Zhuo;Xuehai Qian;Hai Helen Li;Yiran Chen
Linghao Song;Jiachen Mao;Youwei Zhuo;Xuehai Qian;Hai Helen Li;Yiran Chen
中科院分区:
其他
文献类型:
--
作者:
Linghao Song;Jiachen Mao;Youwei Zhuo;Xuehai Qian;Hai Helen Li;Yiran Chen

文献摘要

相似文献

近年来,随着人工智能的兴起,深度神经网络(DNN)已在许多领域得到广泛应用。为了实现高性能和能源效率,学术界和工业界对 DNN 的硬件加速(尤其是推理)进行了深入研究。然而,我们仍然面临两个挑战:大型 DNN 模型和数据集,导致频繁的片外内存访问;以及 DNN 的训练,这在最近的加速器设计中没有得到很好的探索。为了真正为深度和大型模型的训练提供高吞吐量和高能效的加速,与大多数现有架构中考虑的层内细粒度并行性相比,我们不可避免地需要使用多个加速器来探索粗粒度并行性。它提出了寻求加速器之间计算和数据流的最佳组织的关键研究问题。在本文中,受机器学习系统近期工作的启发,我们提出了一种解决方案 HyPar,用于通过一系列 DNN 加速器确定深度神经网络训练的分层并行性。 HyPar 对 DNN 加速器的特征图张量(输入和输出)、内核张量、梯度张量和误差张量进行分区。分区构成了加权层并行性的选择。优化目标是搜索一个分区,使训练完整的 DNN 期间的总通信量最小化。为了解决这个问题,我们提出了一个通信模型来解释通信的来源和数量。然后,我们使用分层动态规划方法来搜索每一层的分区。
With the rise of artificial intelligence in recent years, Deep Neural Networks (DNNs) have been widely used in many domains. To achieve high performance and energy efficiency, hardware acceleration (especially inference) of DNNs is intensively studied both in academia and industry. However, we still face two challenges: large DNN models and datasets, which incur frequent off-chip memory accesses; and the training of DNNs, which is not well-explored in recent accelerator designs. To truly provide high throughput and energy efficient acceleration for the training of deep and large models, we inevitably need to use multiple accelerators to explore the coarse-grain parallelism, compared to the fine-grain parallelism inside a layer considered in most of the existing architectures. It poses the key research question to seek the best organization of computation and dataflow among accelerators. In this paper, inspired by recent work in machine learning systems, we propose a solution HyPar to determine layer-wise parallelism for deep neural network training with an array of DNN accelerators. HyPar partitions the feature map tensors (input and output), the kernel tensors, the gradient tensors, and the error tensors for the DNN accelerators. A partition constitutes the choice of parallelism for weighted layers. The optimization target is to search a partition that minimizes the total communication during training a complete DNN. To solve this problem, we propose a communication model to explain the source and amount of communications. Then, we use a hierarchical layer-wise dynamic programming method to search for the partition for each layer.