Verifiable and Compositional Reinforcement Learning Systems

Verifiable and Compositional Reinforcement Learning Systems
复制标题

DOI:
10.1609/icaps.v32i1.19849
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Cyrus Neary;Christos K. Verginis;Murat Cubuktepe;U. Topcu
Cyrus Neary;Christos K. Verginis;Murat Cubuktepe;U. Topcu
中科院分区:
其他
文献类型:
--
作者:
Cyrus Neary;Christos K. Verginis;Murat Cubuktepe;U. Topcu

文献摘要

被引文献

相似文献

我们提出了一个可验证和组合强化学习(RL)的框架,其中RL子系统的集合被组合起来,每个RL子系统学习完成一个单独的子任务,以实现一个整体任务。该框架由一个高层模型和一个低层子系统集合组成,该模型表示为一个参数马尔可夫决策过程(PMDP),用于规划和分析子系统的组成。通过定义子系统之间的接口,该框架使得能够将任务规范(例如,以至少0.95的概率达到目标状态集)自动分解为各个子任务规范,即,假设满足子系统的进入条件,则以至少某个最小概率实现子系统的退出条件。这进而允许对子系统进行独立的训练和测试;如果每个子系统都学习了满足相应子任务规范的策略,那么它们的组成就保证满足总体任务规范。相反,如果学习的策略不能完全满足子任务规范,我们提出了一种方法,该方法被描述为在pMDP中寻找最优参数集的问题,以自动更新子任务规范以解决所观察到的缺陷。其结果是定义子任务规范并训练子系统以满足这些规范的迭代过程。作为一个额外的好处,这一程序允许在培训期间自动确定总体任务中特别具有挑战性或重要的组成部分,并将重点放在这些部分上。实验结果表明,该框架在离散和连续RL环境下都具有新颖的性能。使用最近的策略优化算法训练RL子系统的集合,以导航迷宫环境的不同部分。然后,将交叉迷宫任务规范分解为子任务规范。如果相应的子系统不能在允许的培训预算内学习令人满意的策略,则自动避免具有挑战性的迷宫部分。不必要的子系统根本没有经过培训。其结果是一个组合式RL系统,它有效地学习以满足任务规范。
We propose a framework for verifiable and compositional reinforcement learning (RL) in which a collection of RL subsystems, each of which learns to accomplish a separate subtask, are composed to achieve an overall task. The framework consists of a high-level model, represented as a parametric Markov decision process (pMDP) which is used to plan and to analyze compositions of subsystems, and of the collection of low-level subsystems themselves. By defining interfaces between the subsystems, the framework enables automatic decompositions of task specifications, e.g., reach a target set of states with a probability of at least 0.95, into individual subtask specifications, i.e. achieve the subsystem's exit conditions with at least some minimum probability, given that its entry conditions are met. This in turn allows for the independent training and testing of the subsystems; if they each learn a policy satisfying the appropriate subtask specification, then their composition is guaranteed to satisfy the overall task specification. Conversely, if the subtask specifications cannot all be satisfied by the learned policies, we present a method, formulated as the problem of finding an optimal set of parameters in the pMDP, to automatically update the subtask specifications to account for the observed shortcomings. The result is an iterative procedure for defining subtask specifications, and for training the subsystems to meet them. As an additional benefit, this procedure allows for particularly challenging or important components of an overall task to be identified automatically, and focused on, during training. Experimental results demonstrate the presented framework's novel capabilities in both discrete and continuous RL settings. A collection of RL subsystems are trained, using proximal policy optimization algorithms, to navigate different portions of a labyrinth environment. A cross-labyrinth task specification is then decomposed into subtask specifications. Challenging portions of the labyrinth are automatically avoided if their corresponding subsystems cannot learn satisfactory policies within allowed training budgets. Unnecessary subsystems are not trained at all. The result is a compositional RL system that efficiently learns to satisfy task specifications.