A Generic Framework for Testing Parallel File Systems

A Generic Framework for Testing Parallel File Systems
复制标题

测试并行文件系统的通用框架

DOI:
--
复制
发表时间:
2016
期刊:
2016 1st Joint International Workshop on Parallel Data Storage and data Intensive Scalable Computing Systems (PDSW-DISCS)
影响因子:
--
通讯作者:
Yong Chen
Yong Chen
中科院分区:
--
文献类型:
--
作者:
Jinrui Cao;Simeng Wang;Dong Dai;Mai Zheng;Yong Chen

文献摘要

被引文献

相似文献

大规模并行文件系统是当今最重要的。然而,尽管它们很重要,但与本地存储系统相比,对其故障恢复能力的研究要少得多。最近对本地存储系统的研究暴露了在故障事件下可能导致数据丢失的各种漏洞,这引发了对构建在本地存储系统上的并行文件系统的关注,提出了一种测试大规模并行文件系统故障处理的通用框架。该框架捕获目标系统的所有存储节点上的所有磁盘I/O命令以模拟真实的故障状态,并检查目标系统是否可以恢复到一致状态而不会导致数据丢失。我们已经为Lustre文件系统构建了一个原型。初步测试结果表明,该框架能够揭示Lustre在不同负载和故障条件下的内部I/O行为,为进一步分析并行文件系统的故障恢复奠定了坚实的基础。
Large-scale parallel file systems are of prime importance today. However, despite of the importance, their failure-recovery capability is much less studied compared with local storage systems. Recent studies on local storage systems have exposed various vulnerabilities that could lead to data loss under failure events, which raise the concern for parallel file systems built on top of them.This paper proposes a generic framework for testing the failure handling of large-scale parallel file systems. The framework captures all disk I/O commands on all storage nodes of the target system to emulate realistic failure states, and checks if the target system can recover to a consistent state without incurring data loss. We have built a prototype for the Lustre file system. Our preliminary results show that the framework is able to uncover the internal I/O behavior of Lustre under different workloads and failure conditions, which provides a solid foundation for further analyzing the failure recovery of parallel file systems.