Small-file access in parallel file systems

Small-file access in parallel file systems
复制标题

DOI:
10.1109/ipdps.2009.5161029
复制
发表时间:
2009-05
期刊:
2009 IEEE International Symposium on Parallel & Distributed Processing
影响因子:
--
通讯作者:
P. Carns;S. Lang;R. Ross;M. Vilayannur;J. Kunkel;T. Ludwig
P. Carns;S. Lang;R. Ross;M. Vilayannur;J. Kunkel;T. Ludwig
中科院分区:
其他
文献类型:
--
作者:
P. Carns;S. Lang;R. Ross;M. Vilayannur;J. Kunkel;T. Ludwig

文献摘要

被引文献

相似文献

当今的计算科学需求已经导致了越来越大的并行计算机,并且存储系统已经发展以满足这些需求。在这种环境中使用的并行文件系统越来越专用于为大型I/O操作提取尽可能高的性能,以牺牲其他潜在的工作负载为代价。虽然一些应用程序已经适应了I/O最佳实践,并且可以在这些系统上获得良好的性能,但许多应用程序的自然I/O模式导致生成许多小文件。这些应用程序不能很好地服务于当前的并行文件系统在非常大的规模。本文介绍了五种技术,优化小文件访问的并行文件系统的超大规模系统。这五种技术都在一个单一的并行文件系统(PVFS)中实现,然后在两个测试平台上进行系统评估。微基准测试和mdtest基准测试用于以前所未有的规模评估优化。我们观察到,与使用16,384个核心的领先计算平台上的基准PVFS配置相比,小文件创建率提高了905%,小文件统计率提高了1,106%,小文件删除率提高了727%。
Today's computational science demands have resulted in ever larger parallel computers, and storage systems have grown to match these demands. Parallel file systems used in this environment are increasingly specialized to extract the highest possible performance for large I/O operations, at the expense of other potential workloads. While some applications have adapted to I/O best practices and can obtain good performance on these systems, the natural I/O patterns of many applications result in generation of many small files. These applications are not well served by current parallel file systems at very large scale. This paper describes five techniques for optimizing small-file access in parallel file systems for very large scale systems. These five techniques are all implemented in a single parallel file system (PVFS) and then systematically assessed on two test platforms. A microbenchmark and the mdtest benchmark are used to evaluate the optimizations at an unprecedented scale. We observe as much as a 905% improvement in small-file create rates, 1,106% improvement in small-file stat rates, and 727% improvement in small-file removal rates, compared to a baseline PVFS configuration on a leadership computing platform using 16,384 cores.