Solving Finite Horizon Decentralized POMDPs by Distributed Reinforcement Learning

Solving Finite Horizon Decentralized POMDPs by Distributed Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
Bikramjit Banerjee;J. Lyle;Landon Kraemer;Rajesh Yellamraju
Bikramjit Banerjee;J. Lyle;Landon Kraemer;Rajesh Yellamraju
中科院分区:
其他
文献类型:
--
作者:
Bikramjit Banerjee;J. Lyle;Landon Kraemer;Rajesh Yellamraju

文献摘要

被引文献

相似文献

分散部分可观测马尔可夫决策过程(Dec-POMDPs)为现实的不确定性多智能体协调问题提供了一个强大的建模技术。流行的解决方案技术是集中的,并假设模型的先验知识。我们提出了一种分布式强化学习方法,代理轮流学习对方的政策的最佳反应。这促进了政策计算问题的分散化,并放松了对问题参数的充分知识的依赖。我们推导了最佳反应学习的样本复杂度与容错性之间的关系。我们的主要贡献是表明,即使是“每叶”样本的复杂性可以指数增长的问题地平线。我们的经验表明,即使样本要求设置低于理论要求,我们的学习方法可以产生(近)在一些基准DecPOMDP问题的最佳政策。我们还提出了一个轻微的修改,经验上似乎显着减少学习时间相对较小的影响,学习政策的质量。
Decentralized partially observable Markov decision processes (Dec-POMDPs) offer a powerful modeling technique for realistic multi-agent coordination problems under uncertainty. Prevalent solution techniques are centralized and assume prior knowledge of the model. We propose a distributed reinforcement learning approach, where agents take turns to learn best responses to each other’s policies. This promotes decentralization of the policy computation problem, and relaxes reliance on the full knowledge of the problem parameters. We derive the relation between the sample complexity of best response learning and error tolerance. Our key contribution is to show that even the“per-leaf”sample complexity could grow exponentially with the problem horizon. We show empirically that even if the sample requirement is set lower than what theory demands, our learning approach can produce (near) optimal policies in some benchmark DecPOMDP problems. We also propose a slight modification that empirically appears to significantly reduce the learning time with relatively little impact on the quality of learned policies.