Service Cost Effective and Reliability Aware Job Scheduling Algorithm on Cloud Computing Systems

Service Cost Effective and Reliability Aware Job Scheduling Algorithm on Cloud Computing Systems
复制标题

DOI:
10.1109/tcc.2021.3137323
复制
发表时间:
2023-04
影响因子:
6.5
通讯作者:
Xiaoyong Tang;Yi Liu;Zeng Zeng-Zeng;B. Veeravalli
Xiaoyong Tang;Yi Liu;Zeng Zeng-Zeng;B. Veeravalli
中科院分区:
计算机科学2区
文献类型:
--
作者:
Xiaoyong Tang;Yi Liu;Zeng Zeng-Zeng;B. Veeravalli

文献摘要

相似文献

如今,越来越多的服务通过云计算系统以按使用付费的模式提供给个人和组织。这种业务服务范式遇到了几个云服务质量(QoS)挑战,例如可靠性、成本和响应时间。提高云服务可靠性的最常见机制是主/备份(PB)容错技术。然而,这种可靠性增强技术不可避免地导致多次复制,这导致高服务成本。在认识到这些挑战,我们首先建立了云计算系统资源管理架构。然后,分析了云服务在虚拟机物理资源上的执行可靠性,并使用支持CUDA(Compute Unified Device Architecture)的并行二维长短期记忆神经网络对云虚拟机的软件故障进行预测。第三,我们提出了一种有效的主/备份云服务成本计算方法。为了克服云服务响应时间的约束,我们将响应时间松弛因子集成到该方法中。第四,我们制定了云服务的可靠性和成本意识的作业调度问题,其目标是最小化云服务的总成本和拒绝率,并提高系统的可靠性。第五,提出了一种启发式贪婪可靠性和成本感知作业调度算法(RCJS)。实验结果表明,该算法在平均服务成本和拒绝率方面明显优于最优冗余VM布局(OPVMP)和MIN-MIN算法.与其他两种算法相比,该算法具有良好的可靠性折衷,适用于高可靠性和低成本要求的云服务。
Nowadays, increasing number of services are provided to individuals and organizations through cloud computing systems in a pay-as-you-use model. This business service paradigm encounters several cloud Quality of Service (QoS) challenges, such as reliability, cost, and response time. The most common mechanism to improve cloud service reliability is a primary/backup (PB) fault-tolerant technique. However, this reliability enhancement technique inevitably results in multiple replications, which lead to high service cost. In recognition of these challenges, we first build a cloud computing systems resources management architecture. Then, we analyze the cloud service execution reliability on the physical resources of a VM and used a CUDA (Compute Unified Device Architecture)-enabled parallel two-dimensional long short-term memory neural network to predict the software faults of a cloud VM. Third, we propose an effective primary/backup cloud service cost calculation approach. To overcome the cloud service response time constraint, we integrate a response time slack factor into this method. Fourth, we formulate the cloud service reliability and cost aware job scheduling problem, which aims at minimizing the total cloud service cost and rejection rate, and improving the system reliability. Fifthly, a heuristic greedy reliability and cost aware job scheduling (RCJS) algorithm is proposed. Finally, a performance evaluation is conducted and the experimental results demonstrate that our proposed RCJS algorithm significantly outperforms optimal redundant VM placement (OPVMP), MIN-MIN algorithms in terms of average service cost and rejection rate. This algorithm also demonstrates good trade-off of reliability when compared to the other two algorithms and is suitable for cloud services with high reliability and low-cost requirements.