Online Performance Modeling and Prediction for Single-VM Applications in Multi-Tenant Clouds

Online Performance Modeling and Prediction for Single-VM Applications in Multi-Tenant Clouds
复制标题

DOI:
10.1109/tcc.2021.3078690
复制
发表时间:
2023-01
影响因子:
6.5
通讯作者:
Hamidreza Moradi;Wen Wang;Dakai Zhu
Hamidreza Moradi;Wen Wang;Dakai Zhu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hamidreza Moradi;Wen Wang;Dakai Zhu

文献摘要

被引文献

相似文献

云计算因其支持灵活的资源需求和低成本而被许多组织广泛采用,这通常通过在多个云租户之间共享底层硬件来实现。然而,这种共享与虚拟机(VM)中的资源争用的变化可能导致云应用的性能的大的变化,这使得普通云用户难以估计其应用的运行时性能。在本文中,我们提出了在线学习方法,用于对在多租户云上重复运行的应用程序进行性能建模和预测(例如在线数据分析任务)。在这里,一些微基准测试被用来探测目标VM的CPU、内存和I/O组件的原位可感知性能。然后,基于这样的剖析信息和现场测量的应用程序的性能,可以用回归或神经网络技术导出预测模型。特别是,为了解决VM的资源竞争强度随时间的变化及其对目标应用程序的影响,我们提出了周期性的模型再训练,其中利用滑动窗口技术来控制用于模型再训练的频率和历史数据。此外,还设计了一种渐进式建模方法,其中回归和神经网络模型逐渐更新,以更好地适应最近的资源争用变化。我们考虑了PARSEC、NAS Parallel和CloudSuite基准测试中的17个代表性应用程序,对所提出的在线方案进行了广泛的评估,以确定最终模型的预测准确性以及私有云和公共云上的相关开销。评估结果表明,即使在资源竞争激烈且变化剧烈的私有云上,所考虑的模型的平均预测误差也可以小于20%。预测误差通常随着更高的再训练频率和更多的历史数据点而减小,但会导致更高的运行时开销。此外,使用神经网络渐进模型,平均预测误差可以减少约7%,同时在私有云上大大减少运行时开销(高达265倍)。对于资源竞争较少的公共云,对于我们提出的在线方案,所考虑的模型的平均预测误差可以小于4%。
Clouds have been adopted widely by many organizations for their supports of flexible resource demands and low cost, which is normally achieved through sharing the underlying hardware among multiple cloud tenants. However, such sharing with the changes in resource contentions in virtual machines (VMs) can result in large variations for the performance of cloud applications, which makes it difficult for ordinary cloud users to estimate the run-time performance of their applications. In this article, we propose online learning methodologies for performance modeling and prediction of applications that run repetitively on multi-tenant clouds (such as on-line data analytic tasks). Here, a few micro-benchmarks are utilized to probe the in-situ perceivable performance of CPU, memory and I/O components of the target VM. Then, based on such profiling information and in-place measured application’s performance, the predictive models can be derived with either Regression or Neural-Network techniques. In particular, to address the changes in the intensity of resource contentions of a VM over time and its effects on the target application, we proposed periodic model retraining where the sliding-window technique was exploited to control the frequency and historical data used for model retraining. Moreover, a progressive modeling approach has been devised where the Regression and Neural-Network models are gradually updated for better adaptation to recent changes in resource contention. With 17 representative applications from PARSEC, NAS Parallel and CloudSuite benchmarks being considered, we have extensively evaluated the proposed online schemes for the prediction accuracy of the resulting models and associated overheads on both a private and public clouds. The evaluation results show that, even on the private cloud with high and radically changed resource contention, the average prediction errors of the considered models can be less than 20 percent with periodic retraining. The prediction errors generally decrease with higher retraining frequencies and more historical data points but incurring higher run-time overheads. Furthermore, with the neural-network progressive models, the average prediction errors can be reduced by about 7 percent with much reduced run-time overheads (up to 265X) on the private cloud. For public clouds with less resource contentions, the average prediction errors can be less than 4 percent for the considered models with our proposed online schemes.