Understanding, Detecting and Localizing Partial Failures in Large System Software

Understanding, Detecting and Localizing Partial Failures in Large System Software
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
Chang Lou;Peng Huang;Scott F. Smith
Chang Lou;Peng Huang;Scott F. Smith
中科院分区:
其他
文献类型:
--
作者:
Chang Lou;Peng Huang;Scott F. Smith

文献摘要

相似文献

部分失败在云系统中经常发生,不幸的是,这些失败的不一致,这些失败并不理解。了解它们的特征的系统。触发将omegagen应用于六个大型分布式系统。在中位检测时间为4.2秒,并确定生成的看门狗的故障范围。
Partial failures occur frequently in cloud systems and can cause serious damage including inconsistency and data loss. Unfortunately, these failures are not well understood. Nor can they be effectively detected. In this paper, we first study 100 real-world partial failures from five mature systems to understand their characteristics. We find that these failures are caused by a variety of defects that require the unique conditions of the production environment to be triggered. Manually writing effective detectors to systematically detect such failures is both time-consuming and error-prone. We thus propose OmegaGen, a static analysis tool that automatically generates customized watchdogs for a given program by using a novel program reduction technique. We have successfully applied OmegaGen to six large distributed systems. In evaluating 22 real-world partial failure cases in these systems, the generated watchdogs can detect 20 cases with a median detection time of 4.2 seconds, and pinpoint the failure scope for 18 cases. The generated watchdogs also expose an unknown, confirmed partial failure bug in the latest version of ZooKeeper.