Decentralized POMDPs

Decentralized POMDPs
复制标题

去中心化 POMDP

DOI:
10.1007/978-3-642-27645-3_15
复制
发表时间:
2012
影响因子:
5.7
通讯作者:
F. Oliehoek
F. Oliehoek
中科院分区:
计算机科学2区
文献类型:
--
作者:
F. Oliehoek

文献摘要

被引文献

相似文献

本章在DEC-POMDP中介绍了分散的POMDP(DEC-POMDP)框架。在执行过程中,代理人的各个策略从历史记录到动作。在最坏的情况下,这一章的时间是在有限的地平线上计划DEC-POMDPS。对于DEC-POMDP的最佳Q值函数,它为其他解决方案方法和其他相关主题提供了指针。
This chapter presents an overview of the decentralized POMDP (Dec-POMDP) framework. In a Dec-POMDP, a team of agents collaborates to maximize a global reward based on local information only. This means that agents do not observe a Markovian signal during execution and therefore the agents’ individual policies map from histories to actions. Searching for an optimal joint policy is an extremely hard problem: it is NEXP-complete. This suggests, assuming NEXP 6=EXP, that any optimal solution method will require doubly exponential time in the worst case. This chapter focuses on planning for Dec-POMDPs over a finite horizon. It covers the forward heuristic search approach to solving Dec-POMDPs, as well as the backward dynamic programming approach. Also, it discusses how these relate to the optimal Q-value function of a Dec-POMDP. Finally, it provides pointers to other solution methods and further related topics.