Deep Coalitional Q-Learning for Dynamic Coalition Formation in Edge Computing

Deep Coalitional Q-Learning for Dynamic Coalition Formation in Edge Computing
复制标题

DOI:
10.1587/transinf.2021kbp0007
复制
发表时间:
2022-05
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Shiyao Ding;Donghui Lin
Shiyao Ding;Donghui Lin
中科院分区:
其他
文献类型:
--
作者:
Shiyao Ding;Donghui Lin

文献摘要

相似文献

随着物联网计算需求的高度发展,资源有限的边缘服务器通常需要协同执行任务。大多数相关研究通常假设静态合作方法,这可能不适合边缘计算的动态环境。在本文中,我们考虑了一种动态合作的方法,即引导边缘服务器动态地形成联盟。它提出了两个问题:1)如何引导它们以最佳方式形成联盟;2)如何处理服务器状态随着任务执行而动态变化的动态特性。本文提出的联合马尔可夫决策过程(CMDP)模型可以很好地处理这些问题。然而,其基本解决方案——联合q -学习,在边缘计算中无法处理任务数量较大时的大规模问题。我们的回应是提出一种称为深度联合q -学习(DCQL)的新算法来解决这个问题。综上所述,我们首先将边缘服务器的动态协作问题表述为一个CMDP:将每个边缘服务器视为一个agent,将动态过程建模为一个MDP, agent观察当前状态形成多个联盟。每个联盟采取一个行动来影响环境,相应地转移到下一个状态来重复上述过程。然后,我们提出了包含深度神经网络的DCQL,可以很好地处理大规模问题。DCQL可以引导边缘服务器以优化某个目标为目标动态地形成联盟。此外,我们运行实验来验证我们提出的算法在不同设置下的有效性。
SUMMARY With the high development of computation requirements in Internet of Things, resource-limited edge servers usually require to cooperate to perform the tasks. Most related studies usually assume a static co-operation approach which might not suit the dynamic environment of edge computing. In this paper, we consider a dynamic cooperation approach by guiding edge servers to form coalitions dynamically. It raises two issues: 1) how to guide them to optimally form coalitions and 2) how to cope with the dynamic feature where server statuses dynamically change as the tasks are performed. The coalitional Markov decision process (CMDP) model proposed in our previous work can handle these issues well. However, its basic solution, coalitional Q-learning, cannot handle the large scale problem when the task number is large in edge computing. Our response is to propose a novel algorithm called deep coalitional Q-learning (DCQL) to solve it. To sum up, we first formulate the dynamic cooperation problem of edge servers as a CMDP: each edge server is regarded as an agent and the dynamic process is modeled as a MDP where the agents observe the current state to formulate several coalitions. Each coalition takes an action to impact the environment which correspondingly transfers to the next state to repeat the above process. Then, we propose DCQL which includes a deep neural network and so can well cope with large scale problem. DCQL can guide the edge servers to form coalitions dynamically with the target of optimizing some goal. Furthermore, we run experiments to verify our proposed algorithm’s e ff ectiveness in di ff erent settings.