Leveraging Sparse Auto-Encoding and Dynamic Learning Rate for Efficient Cloud Workloads Prediction

Leveraging Sparse Auto-Encoding and Dynamic Learning Rate for Efficient Cloud Workloads Prediction
复制标题

DOI:
10.1109/access.2023.3289884
复制
发表时间:
2023
期刊:
影响因子:
3.9
通讯作者:
Dalal Alqahtani
Dalal Alqahtani
中科院分区:
计算机科学3区
文献类型:
--
作者:
Dalal Alqahtani

文献摘要

相似文献

云计算提供了对集中共享计算资源池的简单按需访问。云计算资源的性能和高效利用需要对云工作负载的准确预测。由于云工作负载的动态特性,这是一个具有挑战性的问题,因此很难预测。在这里,我们利用深度学习,通过适当的培训,可以为预测数据中心工作负载提供准确的基础。然而,深度学习(DL)模型的培训具有挑战性。一个挑战是需要定义和调整大量的超参数。通过优化这些超参数可以显著提高神经网络模型的性能。我们认识到使用深度学习高效预测数据中心工作负载的两个基本问题。首先是高维,它需要通过某种形式的降维来去除多余的信息。其次,是学习速度。较低的学习速度会使培训时间过长,而较长的学习速度可能会错过最佳解决方案。因此,我们的方法是双管齐下的。首先,我们使用稀疏自动编码器(SAE)从原始的高维历史云负载数据中检索基本的负载表示。其次,我们使用GRU-SWSLD(GRU-SWSLD)作为学习衰退的门控递归单元。该系统使用来自Google集群工作负载跟踪的数据进行了演示,以使用数据中心在多个连续时间步的工作负载跟踪来预测中央处理器(CPU)的使用情况。我们的实验结果表明,与其他模型相比,我们提出的方法在精度和训练时间之间提供了更好的折衷。
Cloud computing provides simple on-demand access to a centralized shared pool of computing resources. Performance and efficient utilization of cloud computing resources requires accurate prediction of cloud workload. This is a challenging problem due to cloud workloads’ dynamic nature, making it difficult to predict. Here we leverage deep learning which can with proper training provide accurate bases for the prediction of data center workload. Deep Learning (DL) models, however, are challenging to train. One challenge is the vast number of hyperparameters needed to define and tune. The performance of a neural network model can be significantly improved by optimizing these hyperparameters. We recognize two of the essential issues to predict data center workloads using deep learning efficiently. First is the high dimensionality which requires removing superfluous information via some form of dimension reduction. Secondly, is the learning rate. Small learning rates can make the time for training very excessive while long learning rates can miss optimal solutions. Our approach is therefore dual-pronged. First, we use Sparse Auto-Encoder (SAE) to retrieve the essential workloads representations from the original high-dimensional historical cloud workloads data. Secondly, we use Gated Recurrent Unit with a Step-Wise Scheduler for the Learning Decay (GRU-SWSLD). The proposed system is demonstrated with data from Google cluster workload traces to predict Central Processing Unit (CPU) usage using the data center’s workload traces at several consecutive time steps. Our experimental results reveal that our proposed methodology provides a better tradeoff between accuracy and training time when compared with other models.