Training Deep Reactive Policies for Probabilistic Planning Problems

Training Deep Reactive Policies for Probabilistic Planning Problems
复制标题

针对概率规划问题训练深度反应策略

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Automated Planning and Scheduling
影响因子:
--
通讯作者:
Prasad Tadepalli
Prasad Tadepalli
中科院分区:
--
文献类型:
--
作者:
Murugeswari Issakkimuthu;Alan Fern;Prasad Tadepalli

文献摘要

被引文献

相似文献

最先进的概率规划者通常在每一步都应用前瞻搜索和推理来做出决策。虽然这种方法可以实现高质量的决策,但对于需要快速决策的问题来说,它的计算成本可能很高。在本文中,我们研究了深度学习通过快速反应策略取代搜索的潜力。我们专注于Risk中描述的概率规划问题的深度反应策略的监督学习。一个关键的挑战是探索网络架构和训练方法的巨大设计空间,这对之前的深度学习成功至关重要。我们调查了一些选择在这个空间,并进行实验,在一组基准问题。我们的研究结果表明,有效的深度反应策略可以学习许多基准问题,利用规划问题描述来定义网络结构可能是有益的。
State-of-the-art probabilistic planners typically apply look- ahead search and reasoning at each step to make a decision. While this approach can enable high-quality decisions, it can be computationally expensive for problems that require fast decision making. In this paper, we investigate the potential for deep learning to replace search by fast reactive policies. We focus on supervised learning of deep reactive policies for probabilistic planning problems described in RDDL. A key challenge is to explore the large design space of network architectures and training methods, which was critical to prior deep learning successes. We investigate a number of choices in this space and conduct experiments across a set of benchmark problems. Our results show that effective deep reactive policies can be learned for many benchmark problems and that leveraging the planning problem description to define the network structure can be beneficial.