Policy optimization by marginal-map probabilistic inference in generative models
Policy optimization by marginal-map probabilistic inference in generative models
复制标题
生成模型中通过边际图概率推理进行策略优化
DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
P. Poupart
中科院分区:
文献类型:
--
作者:
Igor Kiselev;P. Poupart
While most current work in POMDP planning focus on the development of scalable approximate algorithms, existing techniques often neglect performance guarantees and sacrifice solution quality to improve efficiency. In contrast, our approach to optimizing POMDP controllers by probabilistic inference and obtaining bounded on solution quality can be summarized as follows: (1) re-formulate POMDP planning as a task of marginal-MAP “mix” (max-sum) inference with respect to a new single-DBN generative model, (2) define a dual representation of the MMAP problem and derive a Bayesian variational approximation framework with an upper bound, (3) and design hybrid message-passing algorithms to optimize a POMDP policy by approximate variational MMAP inference in the DBN generative model.