Entropy-Aware I/O Pipelining for Large-Scale Deep Learning on HPC Systems

Entropy-Aware I/O Pipelining for Large-Scale Deep Learning on HPC Systems
复制标题

DOI:
10.1109/mascots.2018.00023
复制
发表时间:
2018-09
期刊:
2018 IEEE 26th International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS)
影响因子:
--
通讯作者:
Yue Zhu;Fahim Chowdhury;Huansong Fu;A. Moody;K. Mohror;Kento Sato;Weikuan Yu
Yue Zhu;Fahim Chowdhury;Huansong Fu;A. Moody;K. Mohror;Kento Sato;Weikuan Yu
中科院分区:
其他
文献类型:
--
作者:
Yue Zhu;Fahim Chowdhury;Huansong Fu;A. Moody;K. Mohror;Kento Sato;Weikuan Yu

文献摘要

被引文献

相似文献

深度神经网络最近因其在计算机视觉和语音识别等广泛应用领域的能力而引起了极大的兴趣。因此,利用领先的高性能计算 (HPC) 系统的前所未有的力量来发挥深度学习的更大潜力非常重要。虽然人们非常重视利用最新的处理器和加速器,但 I/O 支持也需要跟上深度神经网络计算能力的增长。在这项研究中,我们引入了一种名为 DeepIO 的熵感知 I/O 框架,用于 HPC 系统上的大规模深度学习。其总体目标是协调内存、通信和 I/O 资源的使用,以实现数据集的高效训练。 DeepIO 的 I/O 管道采用了多种新颖的优化:RDMA(远程直接内存访问)辅助的原位改组、输入管道和熵感知机会排序。此外,我们还设计了便携式存储接口,以支持任何底层存储系统上的高效 I/O。我们已经将 DeepIO 实现为流行的 TensorFlow 框架的原型,并在各种不同的存储系统上对其进行了评估。我们的评估表明,DeepIO 的性能明显优于现有的基于内存的存储系统。
Deep neural networks have recently gained tremendous interest due to their capabilities in a wide variety of application areas such as computer vision and speech recognition. Thus it is important to exploit the unprecedented power of leadership High-Performance Computing (HPC) systems for greater potential of deep learning. While much attention has been paid to leverage the latest processors and accelerators, I/O support also needs to keep up with the growth of computing power for deep neural networks. In this research, we introduce an entropy-aware I/O framework called DeepIO for large-scale deep learning on HPC systems. Its overarching goal is to coordinate the use of memory, communication, and I/O resources for efficient training of datasets. DeepIO features an I/O pipeline that utilizes several novel optimizations: RDMA (Remote Direct Memory Access)-assisted in-situ shuffling, input pipelining, and entropy-aware opportunistic ordering. In addition, we design a portable storage interface to support efficient I/O on any underlying storage system. We have implemented DeepIO as a prototype for the popular TensorFlow framework and evaluated it on a variety of different storage systems. Our evaluation shows that DeepIO delivers significantly better performance than existing memory-based storage systems.