Multi-Agent Reinforcement Learning for Cooperative Task Offloading in Distributed Edge Cloud Computing

Multi-Agent Reinforcement Learning for Cooperative Task Offloading in Distributed Edge Cloud Computing
复制标题

DOI:
10.1587/transinf.2021dap0010
复制
发表时间:
2022-05
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Shiyao Ding;Donghui Lin
Shiyao Ding;Donghui Lin
中科院分区:
其他
文献类型:
--
作者:
Shiyao Ding;Donghui Lin

文献摘要

相似文献

分布式边缘云计算是物联网的重要计算基础设施,其ffl加载问题近年来备受关注。已有的分布式边缘云计算中ffl加载任务的研究大多假设每个自利的用户都拥有一台边缘服务器,并选择是在本地执行任务,还是将任务ffl到云服务器上。每个边缘服务器的目标是最大化自己的利益,如低延迟成本,这与非合作设置相对应。然而,随着智慧医院、智慧工厂等智能物联网社区的强劲发展,所有边缘和云服务器都可以像科技公司一样属于一个组织。这与组织的目标是最大化团队对整个边缘云计算系统的兴趣的合作设置相对应。在本文中,我们考虑了一个新的问题,称为ffl负载的协作任务,其中所有边缘服务器试图协同工作,以使整个边缘云计算系统达到低延迟代价和低能量代价的良好性能。然而,这个问题很难解决,原因有两个:1)每个边缘服务器的状态动态变化,任务到达不确定;2)每个边缘服务器只能观察到自己的状态,这使得全局信息不可用,很难优化团队兴趣。为了解决这些问题,我们将问题描述为一个分散的部分可观测马尔可夫决策过程(DEC-POMDP),它能够很好地处理部分观测下的动态特征。然后,我们应用了一种称为值分解网络的多ff强化学习算法,并提出了一种基于值分解网络的任务加载算法。具体地说,其动机是使用团队价值函数来评估团队兴趣,然后将团队兴趣划分为每个边缘服务器的个人价值函数。fi。然后,每个边缘服务器在能够最大化团队兴趣的方向上更新其个体价值函数。最后,我们选择了一个真实数据集的一部分来评估我们的算法,结果表明我们的算法与其他一些已有的方法相比是有效的ff。
SUMMARY Distributed edge cloud computing is an important computation infrastructure for Internet of Things (IoT) and its task o ffl oading problem has attracted much attention recently. Most existing work on task o ffl oading in distributed edge cloud computing usually assumes that each self-interested user owns one edge server and chooses whether to execute its tasks locally or to o ffl oad the tasks to cloud servers. The goal of each edge server is to maximize its own interest like low delay cost, which corresponds to a non-cooperative setting. However, with the strong development of smart IoT communities such as smart hospital and smart factory, all edge and cloud servers can belong to one organization like a technology company. This corresponds to a cooperative setting where the goal of the organization is to maximize the team interest in the overall edge cloud computing system. In this paper, we consider a new problem called cooperative task o ffl oading where all edge servers try to cooperate to make the entire edge cloud computing system achieve good performance such as low delay cost and low energy cost. However, this problem is hard to solve due to two issues: 1) each edge server status dynamically changes and task arrival is uncertain; 2) each edge server can observe only its own status, which makes it hard to optimize team interest as global information is unavailable. For solving these issues, we formulate the problem as a decentralized partially observable Markov decision process (Dec-POMDP) which can well handle the dynamic features under partial observations. Then, we apply a multi-agent reinforcement learning algorithm called value decomposition network (VDN) and propose a VDN-based task o ff loading algo-rithm (VDN-TO) to solve the problem. Specifically, the motivation is that we use a team value function to evaluate the team interest, which is then divided into individual value functions for each edge server. Then, each edge server updates its individual value function in the direction that can maximize the team interest. Finally, we choose a part of a real dataset to evaluate our algorithm and the results show the e ff ectiveness of our algorithm in a comparison with some other existing methods.