Learning-Based Resource Allocation in Cloud Data Center using Advantage Actor-Critic

Learning-Based Resource Allocation in Cloud Data Center using Advantage Actor-Critic
复制标题

DOI:
10.1109/icc.2019.8761309
复制
发表时间:
2019-05
期刊:
ICC 2019 - 2019 IEEE International Conference on Communications (ICC)
影响因子:
--
通讯作者:
Zheyi Chen;Jia Hu;G. Min
Zheyi Chen;Jia Hu;G. Min
中科院分区:
其他
文献类型:
--
作者:
Zheyi Chen;Jia Hu;G. Min

文献摘要

相似文献

由于系统状态的不断变化和用户需求的多样性,云数据中心的资源分配面临着动态性和复杂性的巨大挑战。虽然有解决方案,专注于解决这个问题,他们不能有效地响应系统状态和用户需求的动态变化,因为他们依赖于系统的先验知识。因此,如何实现云数据中心资源的自动化和自适应分配,以满足不同的系统需求,仍然是云数据中心面临的一个挑战。为了科普这一挑战,我们提出了一个基于优势行动者-评论家的强化学习(RL)框架,用于云数据中心的资源分配。首先,参与者参数化策略(分配资源),并根据来自批评者的分数(评估动作)选择连续动作(调度作业)。然后,通过梯度上升的方式更新策略,利用优势函数可以显著减小策略梯度的方差。使用Google集群使用跟踪的仿真结果表明了该方法在云资源分配中的有效性。此外,所提出的方法优于经典的资源分配算法的作业延迟,并实现更快的收敛速度比传统的策略梯度方法。
Due to the ever-changing system states and various user demands, resource allocation in cloud data center is faced with great challenges in dynamics and complexity. Although there are solutions that focus on addressing this problem, they cannot effectively respond to the dynamic changes of system states and user demands since they depend on the prior knowledge of the system. Therefore, it is still an open challenge to realize automatic and adaptive resource allocation in order to satisfy diverse system requirements in cloud data center. To cope with this challenge, we propose an advantage actor-critic based reinforcement learning (RL) framework for resource allocation in cloud data center. First, the actor parameterizes the policy (allocating resources) and chooses continuous actions (scheduling jobs) based on the scores (evaluating actions) from the critic. Next, the policy is updated by gradient ascent and the variance of policy gradient can be significantly reduced with the advantage function. Simulations using Google cluster-usage traces show the effectiveness of the proposed method in cloud resource allocation. Moreover, the proposed method outperforms classic resource allocation algorithms in terms of job latency and achieves faster convergence speed than the traditional policy gradient method.