Analyzing and Escaping Local Optima in Planning as Inference for Partially Observable Domains

Analyzing and Escaping Local Optima in Planning as Inference for Partially Observable Domains
复制标题

分析和逃避规划中的局​​部最优作为部分可观察域的推理

DOI:
10.1007/978-3-642-23783-6_39
复制
发表时间:
2011
期刊:
ArXiv
影响因子:
--
通讯作者:
Marc Toussaint
Marc Toussaint
中科院分区:
--
文献类型:
--
作者:
P. Poupart;Tobias Lang;Marc Toussaint

文献摘要

被引文献

相似文献

作为推理的规划是最近出现的一种通用的方法,用于具有离散和连续变量的完全和部分可观测领域中的单主体和多主体系统的决策理论规划和强化学习。由于作为推理的规划本质上解决了当状态部分可观测时的非凸优化问题,因此需要开发能够稳健地逃离局部最优的技术。研究了单智能体部分可观测马尔可夫决策过程(POMDP)中有限状态控制器的局部最优性。我们证明了EM收敛到相对于一步先行最优的控制器。为了避免局部最优,我们提出了两种算法:一种是在控制器中加入节点以确保多步超前搜索的最优性,另一种是以贪婪的方式分割节点以提高奖励概率。这些方法在基准问题上得到了实证验证。
Planning as inference recently emerged as a versatile approach to decision-theoretic planning and reinforcement learning for single and multi-agent systems in fully and partially observable domains with discrete and continuous variables. Since planning as inference essentially tackles a non-convex optimization problem when the states are partially observable, there is a need to develop techniques that can robustly escape local optima. We investigate the local optima of finite state controllers in single agent partially observable Markov decision processes (POMDPs) that are optimized by expectation maximization (EM). We show that EM converges to controllers that are optimal with respect to a one-step lookahead. To escape local optima, we propose two algorithms: the first one adds nodes to the controller to ensure optimality with respect to a multi-step lookahead, while the second one splits nodes in a greedy fashion to improve reward likelihood. The approaches are demonstrated empirically on benchmark problems.