Imitation-Projected Programmatic Reinforcement Learning

Imitation-Projected Programmatic Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
Abhinav Verma;Hoang Minh Le;Yisong Yue;Swarat Chaudhuri
Abhinav Verma;Hoang Minh Le;Yisong Yue;Swarat Chaudhuri
中科院分区:
其他
文献类型:
--
作者:
Abhinav Verma;Hoang Minh Le;Yisong Yue;Swarat Chaudhuri

文献摘要

相似文献

我们研究了程序化强化学习的问题,其中政策用象征性语言表示为简短的程序。与神经政策相比,程序化政策可以更容易解释,可以推广和正式验证;但是,为此类政策设计严格的学习方法仍然是一个挑战。我们应对这一挑战的方法 - 一种称为Propel的元算法 - 基于三个见解。首先,我们将学习任务视为策略领域的优化,模拟所需策略具有程序化表示的约束,并使用镜像下降的形式解决此优化问题,该形式将梯度逐步进入不受限制的策略空间,然后返回。进入约束空间。其次,我们将无约束的政策空间视为混合神经和程序性表示,这使得可以采用最先进的深层政策梯度方法。第三,我们通过模仿学习将投影步骤作为程序综合,并为此任务利用当代组合方法。我们提出了用于推动的理论收敛结果,并在三个连续的控制域中凭经验评估了该方法。实验表明,Propel可以显着胜过学习计划政策的最先进方法。
We study the problem of programmatic reinforcement learning, in which policies are represented as short programs in a symbolic language. Programmatic policies can be more interpretable, generalizable, and amenable to formal verification than neural policies; however, designing rigorous learning approaches for such policies remains a challenge. Our approach to this challenge -- a meta-algorithm called PROPEL -- is based on three insights. First, we view our learning task as optimization in policy space, modulo the constraint that the desired policy has a programmatic representation, and solve this optimization problem using a form of mirror descent that takes a gradient step into the unconstrained policy space and then projects back onto the constrained space. Second, we view the unconstrained policy space as mixing neural and programmatic representations, which enables employing state-of-the-art deep policy gradient approaches. Third, we cast the projection step as program synthesis via imitation learning, and exploit contemporary combinatorial methods for this task. We present theoretical convergence results for PROPEL and empirically evaluate the approach in three continuous control domains. The experiments show that PROPEL can significantly outperform state-of-the-art approaches for learning programmatic policies.