A Coalitional Markov Decision Process Model for Dynamic Coalition Formation among Agents

A Coalitional Markov Decision Process Model for Dynamic Coalition Formation among Agents
复制标题

DOI:
10.1109/wiiat50758.2020.00044
复制
发表时间:
2020-12
期刊:
2020 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT)
影响因子:
--
通讯作者:
Shiyao Ding;Donghui Lin
Shiyao Ding;Donghui Lin
中科院分区:
其他
文献类型:
--
作者:
Shiyao Ding;Donghui Lin

文献摘要

相似文献

在多智能体领域,大多数关于联盟形成问题的研究都假设静态环境,但在现实场景中,联盟形成问题可能发生在动态环境中。这就产生了动态联盟形成问题,其中联盟结构可能会随着时间而变化;我们的回应是提出联合马尔可夫决策过程(CMDP)。在 CMDP 中,动态过程被建模为 MDP,其中代理观察当前状态以制定一些联盟,每个联盟代表一个可以采取行动影响环境的单元。环境有可能转移到下一个状态以重复该过程。然而,在 MDP 转换过程中改变联盟结构会产生一定的成本,阻碍了使用解决 MDP 的经典算法(例如 Q 学习)来解决 CMDP。因此,我们提出了一种新的算法联盟Q学习来解决CMDP,并证明它可以保证CMDP中最优策略的收敛;此外,我们将所提出的算法应用于边缘计算中的动态形成问题,以指导边缘服务器协作执行任务,从而验证算法的有效性。
In multi-agent field, most studies on coalition formation problems assume static environments, but in real-world scenarios coalition formation problems can occur in dynamic environments. This creates the dynamic coalition formation problem where the coalition structure might change with time; our response is to propose a coalitional Markov decision process (CMDP). In CMDP, the dynamic process is modeled as a MDP where the agents observe the current state to formulate some coalitions and each coalition represents a unit that can take action to impact the environment. The environment probabilistically transfers to the next state to repeat the process. However, changing the coalition structure in the midst of MDP transitions incurs a cost that hinders the use of the classical algorithms invoked to solve MDP (e.g. Q-learning) to solve CMDP. Thus, we propose a novel algorithm coalitional Q-learning to solve CMDP and prove it can guarantee the convergence on optimal policies in CMDP; Furthermore, we apply the proposed algorithm to a dynamic formation problem in edge computing to guide edge servers to cooperatively perform tasks and thus verify the algorithm’s effectiveness.