Exploring Flexible Communications for Streamlining DNN Ensemble Training Pipelines

Exploring Flexible Communications for Streamlining DNN Ensemble Training Pipelines
复制标题

DOI:
10.1109/sc.2018.00067
复制
发表时间:
2018-11
期刊:
SC18: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Randall Pittman;Hui Guan;Xipeng Shen;Seung-Hwan Lim;R. Patton
Randall Pittman;Hui Guan;Xipeng Shen;Seung-Hwan Lim;R. Patton
中科院分区:
其他
文献类型:
--
作者:
Randall Pittman;Hui Guan;Xipeng Shen;Seung-Hwan Lim;R. Patton

文献摘要

相似文献

在一组节点上并行训练深度神经网络(DNN)集成是一种常见的做法,即训练多个模型,以构建具有更高预测精度的模型,或快速调整训练模型的参数。现有的集成训练流水线执行大量的冗余操作,导致不必要的CPU使用,甚至是糟糕的流水线性能。为了消除这些冗余,我们需要具有比现有DNN框架所能提供的更大通信灵活性的管道。该项目研究了一系列设计,以提高管道的灵活性和适应性,同时还提高了性能。我们使用TensorFlow和Horovod实现了我们的设计,并使用大型GPU集群-橡树岭国家实验室的泰坦超级计算机-中的几个大型DNN对其进行了测试。结果表明,采用新的灵活的通信机制,训练过程中的CPU时间减少了2-11倍。此外,当施加CPU核心限制时,我们的实施可以实现高达10倍的加速比。与基准相比,我们最好的流程还将合奏训练过程的平均功率消耗降低了5-16%。
Parallel training of a Deep Neural Network (DNN) ensemble on a cluster of nodes is a common practice to train multiple models in order to construct a model with a higher prediction accuracy, or to quickly tune the parameters of a training model. Existing ensemble training pipelines perform a great deal of redundant operations, resulting in unnecessary CPU usage, or even poor pipeline performance. In order to remove these redundancies, we need pipelines with more communication flexibility than existing DNN frameworks can provide. This project investigates a series of designs to improve pipeline flexibility and adaptivity, while also increasing performance. We implement our designs using Tensorflow with Horovod, and test it using several large DNNs in a large scale GPU cluster, the Titan supercomputer at Oak Ridge National Lab. Our results show that with the new flexible communication schemes, the CPU time spent during training is reduced by 2-11X. Furthermore, our implementation can achieve up to 10X speedups when CPU core limits are imposed. Our best pipeline also reduces the average power draw of the ensemble training process by 5-16% when compared to the baseline.