Stochastic Planning with Lifted Symbolic Trajectory Optimization

Stochastic Planning with Lifted Symbolic Trajectory Optimization
复制标题

DOI:
10.1609/icaps.v29i1.3467
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
Hao Cui;Thomas Keller;R. Khardon
Hao Cui;Thomas Keller;R. Khardon
中科院分区:
其他
文献类型:
--
作者:
Hao Cui;Thomas Keller;R. Khardon

文献摘要

相似文献

本文研究了具有大因子状态和动作空间的在线随机规划问题。在最近的工作中,一种很有前途的方法是通过对当前状态下可应用动作的状态进行聚合模拟来估计它们的质量。这导致了显着的加速,相比于搜索具体的状态和动作,并足以指导决策的情况下,随机策略的性能是状态的质量信息。本文对这种方法进行了两个重大改进。第一,灵感来自于提升的信念传播,利用问题的结构,得到一个更紧凑的计算图的聚合模拟。第二个改进取代了随机策略嵌入在计算图中的符号变量,同时优化搜索高质量的行动。这扩大了该方法的范围,以解决需要深度搜索以及信息通过随机步骤快速丢失的问题。实证评估表明,这些想法显着提高性能,导致硬规划问题的最先进的性能。
This paper investigates online stochastic planning for problems with large factored state and action spaces. One promising approach in recent work estimates the quality of applicable actions in the current state through aggregate simulation from the states they reach. This leads to significant speedup, compared to search over concrete states and actions, and suffices to guide decision making in cases where the performance of a random policy is informative of the quality of a state. The paper makes two significant improvements to this approach. The first, taking inspiration from lifted belief propagation, exploits the structure of the problem to derive a more compact computation graph for aggregate simulation. The second improvement replaces the random policy embedded in the computation graph with symbolic variables that are optimized simultaneously with the search for high quality actions. This expands the scope of the approach to problems that require deep search and where information is lost quickly with random steps. An empirical evaluation shows that these ideas significantly improve performance, leading to state of the art performance on hard planning problems.