Decentralized POMDPs
Decentralized POMDPs
复制标题
去中心化 POMDP
DOI:
10.1007/978-3-642-27645-3_15
复制
发表时间:
2012
影响因子:
5.7
通讯作者:
F. Oliehoek
中科院分区:
文献类型:
--
作者:
F. Oliehoek
This chapter presents an overview of the decentralized POMDP (Dec-POMDP) framework. In a Dec-POMDP, a team of agents collaborates to maximize a global reward based on local information only. This means that agents do not observe a Markovian signal during execution and therefore the agents’ individual policies map from histories to actions. Searching for an optimal joint policy is an extremely hard problem: it is NEXP-complete. This suggests, assuming NEXP 6=EXP, that any optimal solution method will require doubly exponential time in the worst case. This chapter focuses on planning for Dec-POMDPs over a finite horizon. It covers the forward heuristic search approach to solving Dec-POMDPs, as well as the backward dynamic programming approach. Also, it discusses how these relate to the optimal Q-value function of a Dec-POMDP. Finally, it provides pointers to other solution methods and further related topics.