PFault: A General Framework for Analyzing the Reliability of High-Performance Parallel File Systems

PFault: A General Framework for Analyzing the Reliability of High-Performance Parallel File Systems
复制标题

PFault:分析高性能并行文件系统可靠性的通用框架

DOI:
10.1145/3205289.3205302
复制
发表时间:
2018
期刊:
Proceedings of the 2018 International Conference on Supercomputing (ICS
影响因子:
--
通讯作者:
Chen, Yong
Chen, Yong
中科院分区:
--
文献类型:
--
作者:
Cao, Jinrui;Gatla, Om Rameshwar;Zheng, Mai;Dai, Dong;Eswarappa, Vidya;Mu, Yan;Chen, Yong

文献摘要

参考文献

被引文献

相似文献

高性能并行文件系统(PFS)是当今最重要的。然而,与本地存储系统相比,其可靠性研究相对较少,这主要是由于缺乏一种有效的分析方法。本文介绍了一个分析PFS故障处理的通用框架。基于一组定义良好的故障模型自动模拟目标PFS中每个存储设备的故障状态,并能够系统地分析故障下PFS的可恢复性。为了验证其实用性,我们应用Pbug对应用最广泛的PFS之一Lustre进行了研究。我们的分析揭示了Lustre的检查和修复实用程序LFSCK失败并出现意外症状(例如,I/O错误、挂起、重新启动)的许多情况。此外,在Pbug的帮助下,我们能够识别资源泄漏问题,即使在运行LFSCK之后,Lustre的内部名称空间和存储空间的一部分也变得不可用。另一方面,我们还验证了最新的Lustre在故障处理方面与以前的版本相比有了明显的改进。我们希望我们的研究和框架能够帮助改进PFSS以实现可靠的高性能计算。
High-performance parallel file systems (PFSes) are of prime importance today. However, despite the importance, their reliability is much less studied compared with that of local storage systems, largely due to the lack of an effective analysis methodology.In this paper, we introduce PFault, a general framework for analyzing the failure handling of PFSes. PFault automatically emulates the failure state of each storage device in the target PFS based on a set of well-defined fault models, and enables analyzing the recoverability of the PFS under faults systematically.To demonstrate the practicality, we apply PFault to study Lustre, one of the most widely used PFSes. Our analysis reveals a number of cases where Lustre's checking and repairing utility LFSCK fails with unexpected symptoms (e.g., I/O error, hang, reboot). Moreover, with the help of PFault, we are able to identify a resource leak problem where a portion of Lustre's internal namespace and storage space become unusable even after running LFSCK. On the other hand, we also verify that the latest Lustre has made noticeable improvement in terms of failure handling comparing to a previous version. We hope our study and framework can help improve PFSes for reliable high-performance computing.
蓝色基因系统应用程序级 I/O 跟踪的早期经验
DOI: --
发表时间: 2008
期刊: 2008 IEEE International Symposium on Parallel and Distributed Processing
影响因子: --
作者:
Seetharami R. Seelam;I. Chung;Ding;H. Wen;Hao Yu
通讯作者: Hao Yu
互联网互联的创新
DOI: --
发表时间: 1988
期刊:
影响因子: --
作者:
C. Partridge
通讯作者: C. Partridge
Blue Gene/P 系统上的应用程序级 I/O 缓存
DOI: --
发表时间: 2009
期刊: 2009 IEEE International Symposium on Parallel & Distributed Processing
影响因子: --
作者:
Seetharami R. Seelam;I. Chung;John Bauer;Hao Yu;H. Wen
通讯作者: H. Wen
测试并行文件系统的通用框架
DOI: --
发表时间: 2016
期刊: 2016 1st Joint International Workshop on Parallel Data Storage and data Intensive Scalable Computing Systems (PDSW-DISCS)
影响因子: --
作者:
Jinrui Cao;Simeng Wang;Dong Dai;Mai Zheng;Yong Chen
通讯作者: Yong Chen