On distributed file tree walk of parallel file systems

On distributed file tree walk of parallel file systems
复制标题

并行文件系统的分布式文件树遍历

DOI:
--
复制
发表时间:
2012
期刊:
International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
John R. Bringhurst
John R. Bringhurst
中科院分区:
--
文献类型:
--
作者:
Jharrod Lafon;S. Misra;John R. Bringhurst

文献摘要

被引文献

相似文献

超级计算机产生大量的数据,通常组织成并行文件系统上的大型目录层次结构。虽然超级计算应用程序是并行的,但用于处理它们的工具(需要完整的目录传输)通常是串行的。我们提出了一个算法框架和三个完全分布式的算法遍历大型并行文件系统,并执行文件操作并行。第一个算法引入了一个随机的工作窃取调度器;第二个改进了第一个接近意识;第三个改进了第二个使用混合方法。我们已经在洛斯阿拉莫斯国家实验室的1.37 petaflop超级计算机Cielo及其7 petabyte文件系统上测试了我们的实现。测试结果表明,我们的算法执行数量级的速度比国家的最先进的算法,同时实现理想的负载平衡和低通信成本。我们目前的性能见解,从我们的算法在生产系统中使用LANL,执行日常文件系统操作。
Supercomputers generate vast amounts of data, typically organized into large directory hierarchies on parallel file systems. While the supercomputing applications are parallel, the tools used to process them requiring complete directory traversais, are typically serial. We present an algorithm framework and three fully distributed algorithms for traversing large parallel file systems, and performing file operations in parallel. The first algorithm introduces a randomized work-stealing scheduler; the second improves the first with proximity-awareness; and the third improves upon the second by using a hybrid approach. We have tested our implementation on Cielo, a 1.37 petaflop supercomputer at the Los Alamos National Laboratory and its 7 petabyte file system. Test results show that our algorithms execute orders of magnitude faster than state-of-the-art algorithms while achieving ideal load balancing and low communication cost. We present performance insights from the use of our algorithms in production systems at LANL, performing daily file system operations.