Auditing the Structural Reliability of the Clouds Ennan

Auditing the Structural Reliability of the Clouds Ennan
复制标题

云恩南结构可靠性审核

DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
B. Ford
B. Ford
中科院分区:
--
文献类型:
--
作者:
Zhai;D. Wolinsky;Hongda Xiao;H. Liu;X. Su;B. Ford

文献摘要

被引文献

相似文献

大规模系统(在云计算中常见)依靠冗余来获得可靠性和可用性。现代云变得越来越复杂和多样化,创造出大型混乱,在发生故障时会遇到长时间的停电。尽管在发生故障后仍存在重大努力,但我们提出了一种新颖的方法,可以通过审核云的基础结构在发生这种混乱的情况下弄清楚这种混乱,我们称之为云结构可靠性审核员(SRA)。 SRA通过以下步骤审核云来实现我们的目标:1)收集综合组件及其依赖性信息,2)使用此数据构造系统范围的故障树,3)和利用故障树分析算法来确定和等级集基于导致云服务中断的可能性的组件。 SRA使云管理员能够事先评估云中的风险,并在重大故障事件发生之前提高其服务部署的可靠性。我们已经建立了执行所有三个任务的原型实现。使用此原型,我们的实验评估表明SRA是实用的:审核包含13,824台服务器的云,而3,000个开关的花费约为6个小时。
Large scale systems, common in cloud computing, rely on redundancy for reliability and availability. Modern clouds have become ever-increasingly complex and diverse creating large messes that experience long outages when failures occur. While there exist significant effort in resolving faults after they occur, we propose a novel approach to untangling this mess before it occurs by auditing the underlying structure of a cloud, which we call the cloud Structural Reliability Auditor (SRA). SRA achieves our goal by auditing a cloud with the following steps: 1) collecting comprehensive component and its dependency information, 2) using this data to construct a system-wide fault tree, 3) and leveraging fault tree analysis algorithms to determine and rank sets of components based on the likelihood of causing a cloud service outage. SRA enables a cloud administrator to be able to evaluate risks within the cloud beforehand and improve the reliability of her service deployments before the occurrences of critical failure events. We have built a prototype implementation that performs all three tasks. Using this prototype, our experimental evaluation shows that SRA is practical: auditing a cloud containing 13,824 servers and 3,000 switches spends about 6 hours.