Programmatic Reinforcement Learning without Oracles

Programmatic Reinforcement Learning without Oracles
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Wenjie Qiu;He Zhu
Wenjie Qiu;He Zhu
中科院分区:
其他
文献类型:
--
作者:
Wenjie Qiu;He Zhu

文献摘要

相似文献

深度强化学习 (RL) 在许多具有挑战性的控制任务中取得了令人鼓舞的成功。然而,由于难以识别模型的控制逻辑与其网络结构的关系,深度强化学习模型缺乏可解释性。以更可解释的表示形式构建的计划性政策成为一种有前途的解决方案。然而仍然存在两个缺点:首先,综合程序化策略需要对程序架构的离散且不可微的搜索空间进行优化。以前的工作并不是最优的,因为它们只枚举了由预训练的强化学习预言机贪婪引导的程序架构。其次,这些工作没有利用组合性(一个重要的编程概念)来重用和组合原始函数以形成新任务的复杂函数。我们的第一个贡献是一个可编程解释的强化学习框架,该框架在由编程语言语法规则定义的架构空间的连续松弛之上进行程序架构搜索。我们的算法允许使用有效的策略梯度方法通过双层优化来学习策略架构和策略参数,因此不需要预先训练的预言机。我们的第二个贡献是通过将学习到的原始函数整合为一个复合程序来解决新的强化学习问题,从而改进程序策略以支持组合性。实验结果表明,我们的算法擅长发现高度可解释的最佳编程策略。
Deep reinforcement learning (RL) has led to encouraging successes in many challenging control tasks. However, a deep RL model lacks interpretability due to the difficulty of identifying how the model’s control logic relates to its network structure. Programmatic policies structured in more interpretable representations emerge as a promising solution. Yet two shortcomings remain: First, synthesizing programmatic policies requires optimizing over the discrete and non-differentiable search space of program architectures. Previous works are suboptimal because they only enumerate program architectures greedily guided by a pretrained RL oracle. Second, these works do not exploit compositionality, an important programming concept, to reuse and compose primitive functions to form a complex function for new tasks. Our first contribution is a programmatically interpretable RL framework that conducts program architecture search on top of a continuous relaxation of the architecture space defined by programming language grammar rules. Our algorithm allows policy architectures to be learned with policy parameters via bilevel optimization using efficient policy-gradient methods, and thus does not require a pretrained oracle. Our second contribution is improving programmatic policies to support compositionality by integrating primitive functions learned to grasp task-agnostic skills as a composite program to solve novel RL problems. Experiment results demonstrate that our algorithm excels in discovering optimal programmatic policies that are highly interpretable.