The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs With Hybrid Parallelism

The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs With Hybrid Parallelism
复制标题

深度学习中强可伸缩性的案例:用混合并行训练大型3D CNN

DOI:
10.1109/tpds.2020.3047974
复制
发表时间:
2020-07
影响因子:
5.3
通讯作者:
Yosuke Oyama;N. Maruyama;Nikoli Dryden;Erin McCarthy;P. Harrington;J. Balewski;S. Matsuoka;Peter Nug
Yosuke Oyama;N. Maruyama;Nikoli Dryden;Erin McCarthy;P. Harrington;J. Balewski;S. Matsuoka;Peter Nug
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yosuke Oyama;N. Maruyama;Nikoli Dryden;Erin McCarthy;P. Harrington;J. Balewski;S. Matsuoka;Peter Nug

文献摘要

相似文献

我们提出了用于训练大规模3D卷积神经网络的可扩展混合并行算法。基于深度学习的新兴科学工作流通常需要使用大型高维样本进行模型训练,这可能会使训练成本更高,甚至由于过度使用内存而不可行。我们通过在整个端到端训练管道(包括计算和I/O)中广泛应用混合并行来解决这些挑战。我们的混合并行算法扩展了标准的数据并行与空间并行,分区在空间域中的一个单一的样本,实现强大的扩展超出小批量尺寸与更大的聚合内存容量。我们用两个具有挑战性的3D CNN,CosmoFlow和3D U-Net来评估我们提出的训练算法。我们的综合性能研究表明,使用高达2K GPU的两个网络都可以实现良好的弱扩展和强扩展。更重要的是,我们能够使用比以前更大的样本来训练CosmoFlow,从而实现预测准确性的数量级提高。
We present scalable hybrid-parallel algorithms for training large-scale 3D convolutional neural networks. Deep learning-based emerging scientific workflows often require model training with large, high-dimensional samples, which can make training much more costly and even infeasible due to excessive memory usage. We solve these challenges by extensively applying hybrid parallelism throughout the end-to-end training pipeline, including both computations and I/O. Our hybrid-parallel algorithm extends the standard data parallelism with spatial parallelism, which partitions a single sample in the spatial domain, realizing strong scaling beyond the mini-batch dimension with a larger aggregated memory capacity. We evaluate our proposed training algorithms with two challenging 3D CNNs, CosmoFlow and 3D U-Net. Our comprehensive performance studies show that good weak and strong scaling can be achieved for both networks using up to 2K GPUs. More importantly, we enable training of CosmoFlow with much larger samples than previously possible, realizing an order-of-magnitude improvement in prediction accuracy.