FanStore: Enabling Efficient and Scalable I/O for Distributed Deep Learning

FanStore: Enabling Efficient and Scalable I/O for Distributed Deep Learning
复制标题

FanStore:为分布式深度学习实现高效且可扩展的 I/O

DOI:
--
复制
发表时间:
2018
期刊:
arXiv.org
影响因子:
--
通讯作者:
N. Gaffney
N. Gaffney
中科院分区:
--
文献类型:
--
作者:
Zhao Zhang;Lei Huang;U. Manor;Linjing Fang;G. Merlo;C. Michoski;J. Cazes;N. Gaffney

文献摘要

参考文献

被引文献

相似文献

新兴的深度学习(DL)应用程序在计算机集群上引入了繁重的I/O工作负载。固有的持久、重复和随机的文件访问模式很容易使元数据和数据服务饱和,并对其他用户产生负面影响。在本文中,我们提出了FanStore,一个短暂的运行时文件系统,优化现有的硬件/软件栈上的DL I/O。FanStore将数据集分发到计算节点的本地存储,并维护全局命名空间。通过系统调用拦截、分布式元数据管理和通用数据压缩等技术,FanStore以高效和可扩展的方式提供了一个符合POSIX标准的接口,并具有本地硬件吞吐量。用户不必进行侵入性的代码更改即可使用FanStore并利用优化的I/O。我们的基准测试和真实的应用程序的实验表明,FanStore可以扩展DL训练到512个计算节点,扩展效率超过90\%。
Emerging Deep Learning (DL) applications introduce heavy I/O workloads on computer clusters. The inherent long lasting, repeated, and random file access pattern can easily saturate the metadata and data service and negatively impact other users. In this paper, we present FanStore, a transient runtime file system that optimizes DL I/O on existing hardware/software stacks. FanStore distributes datasets to the local storage of compute nodes, and maintains a global namespace. With the techniques of system call interception, distributed metadata management, and generic data compression, FanStore provides a POSIX-compliant interface with native hardware throughput in an efficient and scalable manner. Users do not have to make intrusive code changes to use FanStore and take advantage of the optimized I/O. Our experiments with benchmarks and real applications show that FanStore can scale DL training to 512 compute nodes with over 90\% scaling efficiency.
DOI: 10.1021/acs.molpharmaceut.6b00248
发表时间: 2016-07-05
影响因子: 4.9
作者:
Aliper A;Plis S;Artemov A;Ulloa A;Mamoshina P;Zhavoronkov A
通讯作者: Zhavoronkov A