Lobster: Load Balance-Aware I/O for Distributed DNN Training

Lobster: Load Balance-Aware I/O for Distributed DNN Training
复制标题

DOI:
10.1145/3545008.3545090
复制
发表时间:
2022-08
期刊:
Proceedings of the 51st International Conference on Parallel Processing
影响因子:
--
通讯作者:
Jie Liu-;Bogdan Nicolae;Dong Li
Jie Liu-;Bogdan Nicolae;Dong Li
中科院分区:
其他
文献类型:
--
作者:
Jie Liu-;Bogdan Nicolae;Dong Li

文献摘要

相似文献

可以通过在GPU等加速器上进行优化和/或缩放计算来加速培训深度神经网络(DNN)的耗时和耗时的过程。 I/o操作,但我们在本文中提出了一种新的方法,以解决其他方法,我们在本文中提出了一种新的整体方法,该方法无法正确解决其他方法:i/o负载不平衡。由于驱逐的蚀刻训练需要以后需要的样本,我们首先对训练样品进行了研究,因为训练样品通过数据加载和预处理,我们描述了龙虾,这是一种数据加载运行时,该运行时使用性能建模和先进的启发式启发式启发式螺纹管理与优化的固定范围。模型和数据集表明,与最先进的方法相比,龙虾方法将I/O的高架和端到端培训时间降低了1.5倍。
The resource-hungry and time-consuming process of training Deep Neural Networks (DNNs) can be accelerated by optimizing and/or scaling computations on accelerators such as GPUs. However, the loading and pre-processing of training samples then often emerges as a new bottleneck. This data loading process engages a complex pipeline that extends from the sampling of training data on external storage to delivery of those data to GPUs, and that comprises not only expensive I/O operations but also decoding, shuffling, batching, augmentation, and other operations. We propose in this paper a new holistic approach to data loading that addresses three challenges not sufficiently addressed by other methods: I/O load imbalances among the GPUs on a node; rigid resource allocations to data loading and data preprocessing steps, which lead to idle resources and bottlenecks; and limited efficiency of caching strategies based on pre-fetching due to eviction of training samples needed soon at the expense of those needed later. We first present a study of key bottlenecks observed as training samples flow through the data loading and preprocessing pipeline. Then, we describe Lobster, a data loading runtime that uses performance modeling and advanced heuristics to combine flexible thread management with optimized eviction for distributed caching in order to mitigate I/O overheads and load imbalances. Experiments with a range of models and datasets show that the Lobster approach reduces both I/O overheads and end-to-end training times by up to 1.5 × compared with state-of-the-art approaches.