Using Learned Policies in Heuristic-Search Planning

Using Learned Policies in Heuristic-Search Planning
复制标题

在启发式搜索规划中使用学习策略

DOI:
--
复制
发表时间:
2007
期刊:
International Joint Conference on Artificial Intelligence
影响因子:
--
通讯作者:
R. Givan
R. Givan
中科院分区:
--
文献类型:
--
作者:
S. Yoon;Alan Fern;R. Givan

文献摘要

被引文献

相似文献

许多当前国家的最先进的规划依赖于前向启发式搜索。这种搜索的成功通常取决于启发式的距离到目标的估计来自plangeland。这样的估计是有效的,在许多领域的指导搜索,但仍有许多其他领域,目前的算法是不够的,以指导有效地向前搜索。在其中的一些领域中,可以从解决许多问题的示例计划中学习反应式策略。然而,由于这些学习技术的归纳性质,政策往往是错误的,并未能实现高成功率。在这项工作中,我们考虑如何有效地整合不完善的学习政策与不完善的策略,以改善每一个单独的。我们提出了一个简单的方法,使用的政策,以增加在每个搜索步骤中扩展的状态。特别地,在每个搜索节点扩展期间,我们不仅添加它的邻居,而且添加从节点直到某个水平线的策略所遵循的轨迹上的所有节点沿着。实证结果表明,我们提出的方法有利于利用自动化技术、学习和启发式搜索,在大多数基准规划领域中表现优于最先进的技术。
Many current state-of-the-art planners rely on forward heuristic search. The success of such search typically depends on heuristic distance-to-the-goal estimates derived from the plangraph. Such estimates are effective in guiding search for many domains, but there remain many other domains where current heuristics are inadequate to guide forward search effectively. In some of these domains, it is possible to learn reactive policies from example plans that solve many problems. However, due to the inductive nature of these learning techniques, the policies are often faulty, and fail to achieve high success rates. In this work, we consider how to effectively integrate imperfect learned policies with imperfect heuristics in order to improve over each alone. We propose a simple approach that uses the policy to augment the states expanded during each search step. In particular, during each search node expansion, we add not only its neighbors, but all the nodes along the trajectory followed by the policy from the node until some horizon. Empirical results show that our proposed approach benefits both of the leveraged automated techniques, learning and heuristic search, outperforming the state-of-the-art in most benchmark planning domains.