Why Globally Re-shuffle? Revisiting Data Shuffling in Large Scale Deep Learning

Why Globally Re-shuffle? Revisiting Data Shuffling in Large Scale Deep Learning
复制标题

DOI:
10.1109/ipdps53621.2022.00109
复制
发表时间:
2022-05
期刊:
2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
Thao Nguyen;François Trahay;Jens Domke;Aleksandr Drozd;Emil;Vatai;Jianwei Liao;M. Wahib;
Thao Nguyen;François Trahay;Jens Domke;Aleksandr Drozd;Emil;Vatai;Jianwei Liao;M. Wahib;
中科院分区:
其他
文献类型:
--
作者:
Thao Nguyen;François Trahay;Jens Domke;Aleksandr Drozd;Emil;Vatai;Jianwei Liao;M. Wahib;

文献摘要

相似文献

随机梯度下降(SGD)是训练深神经网络(DNN)的最普遍的算法。 SGD以随机访问方式迭代每个训练时期处理数据样本中的输入数据集。因为这给I/O子系统施加了巨大的压力,所以在HPC环境中分布SGD的最常见方法是将整个数据集复制到节点local SSD。但是,由于数据集迅速增长,这种方法变得越来越不可行。令人惊讶的是,从经验的角度来看,在文献中没有得到很多关注的问题以及在何种程度上需要随机访问的问题。在本文中,我们重新审视DL工作负载中的数据改组,以调查工人之间数据集分区并仅在每个培训时期内进行部分分布式交换的可行性。通过对高达2,048 GPU的ABCI和4,096个Fugaku的计算节点的广泛实验,我们证明,在仔细调整部分分布式交换时,可以保持全球改组的验证精度。我们提供了Pytorch中实现的解决方案,使用户能够控制提出的数据交换方案。
Stochastic gradient descent (SGD) is the most prevalent algorithm for training Deep Neural Networks (DNN). SGD iterates the input data set in each training epoch processing data samples in a random access fashion. Because this puts enormous pressure on the I/O subsystem, the most common approach to distributed SGD in HPC environments is to replicate the entire dataset to node local SSDs. However, due to rapidly growing data set sizes this approach has become increasingly infeasible. Surprisingly, the questions of why and to what extent random access is required have not received a lot of attention in the literature from an empirical standpoint. In this paper, we revisit data shuffling in DL workloads to investigate the viability of partitioning the dataset among workers and performing only a partial distributed exchange of samples in each training epoch. Through extensive experiments on up to 2,048 GPUs of ABCI and 4,096 compute nodes of Fugaku, we demonstrate that in practice validation accuracy of global shuffling can be maintained when carefully tuning the partial distributed exchange. We provide a solution implemented in PyTorch that enables users to control the proposed data exchange scheme.