Neurosymbolic Reinforcement Learning with Formally Verified Exploration

Neurosymbolic Reinforcement Learning with Formally Verified Exploration
复制标题

DOI:
--
复制
发表时间:
2020-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Greg Anderson;Abhinav Verma;Işıl Dillig;Swarat Chaudhuri
Greg Anderson;Abhinav Verma;Işıl Dillig;Swarat Chaudhuri
中科院分区:
其他
文献类型:
--
作者:
Greg Anderson;Abhinav Verma;Işıl Dillig;Swarat Chaudhuri

文献摘要

被引文献

相似文献

我们提出了Revel,一个部分神经强化学习(RL)框架,用于在连续状态和动作空间中进行可证明安全的探索。可证明安全的深度RL的一个关键挑战是,在学习循环中重复验证神经网络在计算上是不可行的。我们使用两个策略类来解决这一挑战:一个是具有近似梯度的一般神经符号类,另一个是允许有效验证的更受限制的符号策略类。我们的学习算法是对策略的镜像下降:在每次迭代中,它安全地将符号策略提升到神经符号空间中,对生成的策略执行安全梯度更新,并将更新后的策略投影到安全的符号子集中,所有这些都不需要神经网络的显式验证。我们的实证结果表明,Revel在许多情况下强制执行安全探索,其中约束策略优化没有,并且它可以发现优于通过先前的方法来验证探索的策略。
We present Revel, a partially neural reinforcement learning (RL) framework for provably safe exploration in continuous state and action spaces. A key challenge for provably safe deep RL is that repeatedly verifying neural networks within a learning loop is computationally infeasible. We address this challenge using two policy classes: a general, neurosymbolic class with approximate gradients and a more restricted class of symbolic policies that allows efficient verification. Our learning algorithm is a mirror descent over policies: in each iteration, it safely lifts a symbolic policy into the neurosymbolic space, performs safe gradient updates to the resulting policy, and projects the updated policy into the safe symbolic subset, all without requiring explicit verification of neural networks. Our empirical results show that Revel enforces safe exploration in many scenarios in which Constrained Policy Optimization does not, and that it can discover policies that outperform those learned through prior approaches to verified exploration.