Solving Finite Horizon Decentralized POMDPs by Distributed Reinforcement Learning
Solving Finite Horizon Decentralized POMDPs by Distributed Reinforcement Learning
复制标题
DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Bikramjit Banerjee;J. Lyle;Landon Kraemer;Rajesh Yellamraju
中科院分区:
文献类型:
--
作者:
Bikramjit Banerjee;J. Lyle;Landon Kraemer;Rajesh Yellamraju
Decentralized partially observable Markov decision processes (Dec-POMDPs) offer a powerful modeling technique for realistic multi-agent coordination problems under uncertainty. Prevalent solution techniques are centralized and assume prior knowledge of the model. We propose a distributed reinforcement learning approach, where agents take turns to learn best responses to each other’s policies. This promotes decentralization of the policy computation problem, and relaxes reliance on the full knowledge of the problem parameters. We derive the relation between the sample complexity of best response learning and error tolerance. Our key contribution is to show that even the“per-leaf”sample complexity could grow exponentially with the problem horizon. We show empirically that even if the sample requirement is set lower than what theory demands, our learning approach can produce (near) optimal policies in some benchmark DecPOMDP problems. We also propose a slight modification that empirically appears to significantly reduce the learning time with relatively little impact on the quality of learned policies.