Synthesizing Programmatic Policies that Inductively Generalize

Synthesizing Programmatic Policies that Inductively Generalize
复制标题

DOI:
--
复制
发表时间:
2020-04
期刊:
ArXiv
影响因子:
--
通讯作者:
J. Inala;O. Bastani;Zenna Tavares;Armando Solar-Lezama
J. Inala;O. Bastani;Zenna Tavares;Armando Solar-Lezama
中科院分区:
其他
文献类型:
--
作者:
J. Inala;O. Bastani;Zenna Tavares;Armando Solar-Lezama

文献摘要

相似文献

深度强化学习已经成功地解决了许多具有挑战性的控制任务。然而,学习策略通常很难推广到新环境。我们提出了一种学习能够捕获重复行为的程序化状态机策略的算法。通过这样做,它们有能力将其推广到需要任意重复次数的实例,这一性质我们称为归纳推广。然而,状态机策略很难学习,因为它们由连续和离散结构组成。我们提出了一种称为自适应教学的学习框架,它通过模仿教师来学习状态机策略;与传统的模仿学习不同,我们的教师根据学生的结构自适应地更新自己。我们展示了如何使用我们的算法来学习归纳概括到新环境的策略,而传统的神经网络策略无法做到这一点。
Deep reinforcement learning has successfully solved a number of challenging control tasks. However, learned policies typically have difficulty generalizing to novel environments. We propose an algorithm for learning programmatic state machine policies that can capture repeating behaviors. By doing so, they have the ability to generalize to instances requiring an arbitrary number of repetitions, a property we call inductive generalization. However, state machine policies are hard to learn since they consist of a combination of continuous and discrete structure. We propose a learning framework called adaptive teaching, which learns a state machine policy by imitating a teacher; in contrast to traditional imitation learning, our teacher adaptively updates itself based on the structure of the student. We show how our algorithm can be used to learn policies that inductively generalize to novel environments, whereas traditional neural network policies fail to do so.