Macro-Action-Based Deep Multi-Agent Reinforcement Learning

Macro-Action-Based Deep Multi-Agent Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2020-04
期刊:
--
影响因子:
--
通讯作者:
Yuchen Xiao;Joshua Hoffman;Chris Amato
Yuchen Xiao;Joshua Hoffman;Chris Amato
中科院分区:
其他
文献类型:
--
作者:
Yuchen Xiao;Joshua Hoffman;Chris Amato

文献摘要

相似文献

在现实世界的多机器人系统中,执行高质量的协作行为需要机器人在不同的时间段内异步推理高级动作选择。宏观行动分散部分可观察马尔可夫决策过程(MacDec-POMDP)为完全协作多智能体任务中的不确定性下的异步决策提供了通用框架。然而,多智能体深度强化学习方法仅针对(同步)原始动作问题而开发。本文提出了两种基于深度 Q 网络(DQN)的方法,用于学习分散式和集中式宏观动作价值函数,并为每种情况引入了新颖的宏观动作轨迹重放缓冲区。对基准问题和更大领域的评估证明了使用宏观操作学习相对于原始操作的优势以及我们方法的可扩展性。
In real-world multi-robot systems, performing high-quality, collaborative behaviors requires robots to asynchronously reason about high-level action selection at varying time durations. Macro-Action Decentralized Partially Observable Markov Decision Processes (MacDec-POMDPs) provide a general framework for asynchronous decision making under uncertainty in fully cooperative multi-agent tasks. However, multi-agent deep reinforcement learning methods have only been developed for (synchronous) primitive-action problems. This paper proposes two Deep Q-Network (DQN) based methods for learning decentralized and centralized macro-action-value functions with novel macro-action trajectory replay buffers introduced for each case. Evaluations on benchmark problems and a larger domain demonstrate the advantage of learning with macro-actions over primitive-actions and the scalability of our approaches.