Abstracting Teammate With Multi-Agent DDPG Using Dual Encoders

Abstracting Teammate With Multi-Agent DDPG Using Dual Encoders
复制标题

DOI:
10.1109/gcce56475.2022.10014178
复制
发表时间:
2022-10
期刊:
2022 IEEE 11th Global Conference on Consumer Electronics (GCCE)
影响因子:
--
通讯作者:
Yuki Hyodo;Takuto Sakuma;Shohei Kato
Yuki Hyodo;Takuto Sakuma;Shohei Kato
中科院分区:
其他
文献类型:
--
作者:
Yuki Hyodo;Takuto Sakuma;Shohei Kato

文献摘要

相似文献

近年来,多智能体强化学习在各个领域得到了广泛的应用,各种学习模型也相继提出。最常见的方法之一是使用所有代理的观察和行动来估计Q值,例如集中式批评,作为对抗非平稳性的对策。虽然这种类型的Q值估计是有效的,但它有一个可伸缩性问题,即随着代理数量的增加,要考虑的信息量也会增加。这个问题使学习变得困难。在本文中,我们提出了一个学习模型,通过将队友的观察和行动的值输入到双编码器中,并提取特征来选择哪个代理的信息对学习有效,从而促进学习。通过与多智能体强化学习中的著名模型MADDPG的比较,验证了该模型的有效性。
In recent years, multi-agent reinforcement learning has been applied in various fields, and learning models have been proposed using a wide variety of approaches. One of the most common is to estimate Q-values using the observations and actions of all agents, such as centralized critic, as a countermeasure against non-stationarity. Although this type of Q-value estimation is effective, it has a scalability problem, which is increasing the amount of information to be considered, as the number of agents increases. The problem makes learning difficult. In this paper, we propose a learning model that facilitates learning by inputting the values of teammates’ observations and actions into dual encoders and extracting features that select which agent’s information is effective for learning. We confirm the effectiveness of the proposed model by comparing with MADDPG, which is a well-known model in multi-agent reinforcement learning.