Proactive Fault-Tolerance Technique to Enhance Reliability of Cloud Service in Cloud Federation Environment

Proactive Fault-Tolerance Technique to Enhance Reliability of Cloud Service in Cloud Federation Environment
复制标题

云联邦环境下增强云服务可靠性的主动容错技术

DOI:
--
复制
发表时间:
2020
影响因子:
6.5
通讯作者:
Sarbani Roy
Sarbani Roy
中科院分区:
计算机科学2区
文献类型:
--
作者:
Benay Kumar Ray;Avirup Saha;Sunirmal Khatua;Sarbani Roy

文献摘要

被引文献

相似文献

云联合是一种新的计算范式,它为云服务提供商(CSP)在其资源需求较低时向其他CSP提供其未使用的资源(虚拟机)铺平了道路。联合还允许CSP在其计算资源需求较高时将其资源请求外包给其他CSP。因此,在云联合环境中,由于CSP能够在它们之间共享资源,服务提供商所提供服务的可靠性和可用性得以提高。此外,为了维持通过联合提供的云服务的可靠性和可用性,联合内成员CSP的计算环境具有容错能力是很重要的。因此,需要一个容错系统来保证云联合环境中云服务的可靠性和可用性。在本文中,我们提出了一种主动容错系统,该系统根据CPU温度来预防联合内的故障。联合内的容错系统被建模为一个多目标优化问题,即在将资源(虚拟机)从有故障的CSP重新分配到联合内无故障的CSP时,最大化利润并最小化迁移成本。为了解决这个问题,我们还提出了一种称为基于偏好的故障管理(PBFM)的算法,以便在发生故障时管理联合。我们进行了大量实验来评估我们所提出机制的有效性,并将其与其他两种机制——迁移成本有保证的故障管理(MCAFM)和利润有保证的故障管理(PAFM)进行比较。结果表明,我们提出的机制PBFM为存在有故障的CSP时利润和迁移成本权衡的一般问题提供了一种优化解决方案。
Cloud federation is a new computing paradigm that has paved the way for cloud service providers (CSPs) to offer their unused resources (virtual machine) to other CSPs when their resource demands are low. Federation also allows CSPs to outsource their resource requests to other CSPs when their computing resources’ demands are high. Thus, in cloud federation environment reliability and availability of services offered by service providers increase as the CSPs are able to share their resources among themselves. Moreover, to maintain the reliability and availability of cloud services offered through federation, it is important that the computational environment of member CSPs within the federation is fault tolerant. Therefore, there is a need for fault tolerant system to guarantee cloud service reliability and availability in cloud federation environment. In this article, we propose a proactive fault tolerance system that preempts faults within the federation on the basis of CPU temperature. The fault tolerance system within the federation is modeled as a multi-objective optimization problem of maximizing profit and minimizing migration cost while redistributing resources (virtual machine) from faulty CSPs to non-faulty CSPs within the federation. To address this issue, we have also proposed an algorithm called Preference Based Fault Management (PBFM) to manage the federation in the event of faults. We perform extensive experiments to evaluate the effectiveness of our proposed mechanism and compare it with two other mechanisms MCAFM (Migration Cost Assured Fault Management) and PAFM (Profit Assured Fault Management). Results show that our proposed mechanism PBFM yields an optimized solution to the general problem of profit and migration cost trade-off in presence of faulty CSPs.