Online generation and use of macro-actions in forward-chaining planning

Online generation and use of macro-actions in forward-chaining planning
复制标题

前向链规划中宏观行动的在线生成和使用

DOI:
--
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
A. Coles
A. Coles
中科院分区:
--
文献类型:
--
作者:
A. Coles

文献摘要

被引文献

相似文献

本文提出了一种在前向链接规划中在线学习和管理宏观动作的技术。宏观动作是在高原(启发式无法提供良好搜索指导的搜索景观区域)上学习的,并在遇到未来高原时重用。生成宏操作库,存储宏操作以用于将来的问题。考虑了管理和修剪此类库的几种策略,允许规划者维护较小的宏观操作集合,以尽量减少对性能的潜在负面影响。这项工作被扩展到研究在搜索过程中不增加分支因子的情况下模拟宏观动作的潜力。在强制爬山 (EHC) 搜索和最佳优先搜索期间,根据在过去的解决方案中当前处于计划尾部的操作执行的次数,对操作进行重新排序。在 EHC 中,这相当于建议两步宏操作而不增加分支因子。这些技术在宽松的规划图启发式下,在具有不同属性的广泛领域进行评估。结果表明,仅使用在线学习技术来选择要保留的最佳宏观操作,摆脱高原宏观操作的库就可以提高规划器的性能。搜索时间修剪被证明是非常有效的;同时根据每个操作的使用次数修剪宏操作库可以提供进一步的改进。行动重新排序已被证明可以提供性能改进,尽管不如使用宏操作所产生的效果那么好,但在通常与宏操作相关的一些问题上不会偶尔出现性能下降。
This thesis presents a technique for online learning and management of macroactions in forward chaining planning. Macro-actions are learnt on plateaux, areas of the search landscape where the heuristic cannot offer good search guidance, and are reused when future plateaux are encountered. Libraries of macro-actions are generated, storing macro-actions for use on future problems. Several strategies for the management and pruning of such libraries are considered allowing the planner to maintain a smaller collection of macro-actions to minimise potential negative effects on performance. The work is extended to investigate the potential for the simulation of macroactions without increasing the branching factor during search. Actions are reordered during enforced hill-climbing (EHC) search, and best-first search, based on the number of times they have followed the action currently at the tail of the plan in past solutions. In EHC this is equivalent to suggesting two-step macroactions without increasing the branching factor. The techniques are evaluated across a wide range of domains with differing properties under the relaxed planning graph heuristic. The results show that a library of plateau-escaping macro-actions can improve planner performance using only online learning techniques to select the best macro-actions to keep. Search time pruning is shown to be very effective; whilst pruning the library of macro-actions based on the number of times each action is used can offer further improvements. Action reordering has been shown to offer performance improvements, albeit not as great as those produced by using macro-actions, but without the occasional degradation in performance on some problems that is often associated with macro-actions.