Policy optimization by marginal-map probabilistic inference in generative models

Policy optimization by marginal-map probabilistic inference in generative models
复制标题

生成模型中通过边际图概率推理进行策略优化

DOI:
--
复制
发表时间:
2014
期刊:
Adaptive Agents and Multi-Agent Systems
影响因子:
--
通讯作者:
P. Poupart
P. Poupart
中科院分区:
--
文献类型:
--
作者:
Igor Kiselev;P. Poupart

文献摘要

被引文献

相似文献

虽然目前POMDP规划的大部分工作都集中在可伸缩近似算法的开发上,但现有的技术往往忽视了性能保证,为了提高效率而牺牲了解的质量。相反,我们通过概率推理优化POMDP控制器和获得解质量有界的方法可以总结如下:(1)将POMDP规划重新描述为关于新的单一DBN生成模型的边际映射混合(max-sum)推理任务;(2)定义MMAP问题的对偶表示,并推导出具有上界的贝叶斯变分近似框架;(3)设计混合消息传递算法,通过在DBN生成模型中的近似变分MMAP推理来优化POMDP策略。
While most current work in POMDP planning focus on the development of scalable approximate algorithms, existing techniques often neglect performance guarantees and sacrifice solution quality to improve efficiency. In contrast, our approach to optimizing POMDP controllers by probabilistic inference and obtaining bounded on solution quality can be summarized as follows: (1) re-formulate POMDP planning as a task of marginal-MAP “mix” (max-sum) inference with respect to a new single-DBN generative model, (2) define a dual representation of the MMAP problem and derive a Bayesian variational approximation framework with an upper bound, (3) and design hybrid message-passing algorithms to optimize a POMDP policy by approximate variational MMAP inference in the DBN generative model.