Data-Driven Robust Multi-Agent Reinforcement Learning

Data-Driven Robust Multi-Agent Reinforcement Learning
复制标题

DOI:
10.1109/mlsp55214.2022.9943500
复制
发表时间:
2022-08
期刊:
2022 IEEE 32nd International Workshop on Machine Learning for Signal Processing (MLSP)
影响因子:
--
通讯作者:
Yudan Wang;Yue Wang;Yi Zhou;Alvaro Velasquez;Shaofeng Zou
Yudan Wang;Yue Wang;Yi Zhou;Alvaro Velasquez;Shaofeng Zou
中科院分区:
其他
文献类型:
--
作者:
Yudan Wang;Yue Wang;Yi Zhou;Alvaro Velasquez;Shaofeng Zou

文献摘要

相似文献

协作环境中的多智能体强化学习(MARL)旨在找到一种联合策略,最大化所有智能体的平均累积奖励。在本文中,我们关注模型不确定性下的 MARL,其中假设转移核位于不确定性集合中,目标是优化不确定性集合上的最坏情况性能。我们研究了无模型设置,其中不确定集围绕未知的马尔可夫决策过程,从中可以顺序获得单个样本轨迹。我们开发了一种强大的多智能体 Q 学习算法,该算法是无模型且完全去中心化的。我们从理论上证明了所提出的算法收敛于极小极大鲁棒策略,并进一步表征了其样本复杂度。与普通的多智能体 Q 学习相比,我们的算法在模型不确定性下提供了可证明的鲁棒性,而不会产生额外的计算和内存成本。
Multi-agent reinforcement learning (MARL) in the collaborative setting aims to find a joint policy that maximizes the accumulated reward averaged over all the agents. In this paper, we focus on MARL under model uncertainty, where the transition kernel is assumed to be in an uncertainty set, and the goal is to optimize the worst-case performance over the uncertainty set. We investigate the model-free setting, where the uncertain set centers around an unknown Markov decision process from which a single sample trajectory can be obtained sequentially. We develop a robust multi-agent Q-learning algorithm, which is model-free and fully decentralized. We theoretically prove that the proposed algorithm converges to the minimax robust policy, and further characterize its sample complexity. Our algorithm, comparing to the vanilla multi-agent Q-learning, offers provable robustness under model uncertainty without incurring additional computational and memory cost.