State Augmentation via Self-Supervision in Offline Multiagent Reinforcement Learning

State Augmentation via Self-Supervision in Offline Multiagent Reinforcement Learning
复制标题

DOI:
10.1109/tcds.2023.3326297
复制
发表时间:
2024-06
影响因子:
5
通讯作者:
Siying Wang;Xiaodie Li;Hong Qu;Wenyu Chen
Siying Wang;Xiaodie Li;Hong Qu;Wenyu Chen
中科院分区:
计算机科学3区
文献类型:
--
作者:
Siying Wang;Xiaodie Li;Hong Qu;Wenyu Chen

文献摘要

相似文献

利用预先收集的离线数据集在没有环境交互的情况下进行学习,使强化学习(RL)在现实世界中取得了重大进展。这种方法对于多智能体RL(MARL)任务也很有吸引力,因为智能体和环境之间存在复杂的交互。然而,与单智能体方法相比,离线MARL面临着更多的挑战,由于更大的状态和动作空间,特别是关于穷人的分布泛化到环境。本研究表明,直接将保守的离线RL算法从单智能体环境转移到多智能体环境是无效的,这是由于累积的外推误差与智能体数量成比例增加。在这篇文章中,我们探讨了三种类型的数据增强技术,可以应用到状态表示的MARL的上下文中的效果。通过结合所提出的数据增强技术与一个国家的最先进的离线多智能体算法,我们提高了集中式$Q $-网络的功能近似。在星际争霸II上进行的实验结果有力地支持了数据增强技术在增强状态空间中离线MARL性能方面的有效性。
The utilization of precollected offline data sets for learning in the absence of environmental interaction has enabled reinforcement learning (RL) to make significant strides in real-world circumstances. This approach is also attractive for multiagent RL (MARL) tasks, given the complex interactions that occur between agents and the environment. However, when compared to the single-agent approach, offline MARL faces more challenges due to the larger state and action space, particularly with regard to poor out-of-distribution generalization to the environment. The present study demonstrates the ineffectiveness of directly transferring conservative offline RL algorithms from single-agent settings to multiagent environments, which is due to the accumulating extrapolation errors that increase in proportion to the number of agents. In this article, we explore the efficacy of three types of data augmentation techniques that can be applied to the state representation in the context of MARL. By combining the proposed data augmentation techniques with a state-of-the-art offline multiagent algorithm, we improve the function approximation of centralized $Q$ -networks. The experimental results conducted on StarCraft II strongly support the effectiveness of the data augmentation techniques in enhancing the performance of offline MARL in the state space.