FanStore: Enabling Efficient and Scalable I/O for Distributed Deep Learning
FanStore: Enabling Efficient and Scalable I/O for Distributed Deep Learning
复制标题
FanStore:为分布式深度学习实现高效且可扩展的 I/O
DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
N. Gaffney
中科院分区:
文献类型:
--
作者:
Zhao Zhang;Lei Huang;U. Manor;Linjing Fang;G. Merlo;C. Michoski;J. Cazes;N. Gaffney
Emerging Deep Learning (DL) applications introduce heavy I/O workloads on computer clusters. The inherent long lasting, repeated, and random file access pattern can easily saturate the metadata and data service and negatively impact other users. In this paper, we present FanStore, a transient runtime file system that optimizes DL I/O on existing hardware/software stacks. FanStore distributes datasets to the local storage of compute nodes, and maintains a global namespace. With the techniques of system call interception, distributed metadata management, and generic data compression, FanStore provides a POSIX-compliant interface with native hardware throughput in an efficient and scalable manner. Users do not have to make intrusive code changes to use FanStore and take advantage of the optimized I/O. Our experiments with benchmarks and real applications show that FanStore can scale DL training to 512 compute nodes with over 90\% scaling efficiency.
影响因子:
4.9
作者:
Aliper A;Plis S;Artemov A;Ulloa A;Mamoshina P;Zhavoronkov A
通讯作者:
Zhavoronkov A