LTL-Based Non-Markovian Inverse Reinforcement Learning

LTL-Based Non-Markovian Inverse Reinforcement Learning
复制标题

DOI:
10.5555/3545946.3599102
复制
发表时间:
2021-10
期刊:
--
影响因子:
--
通讯作者:
Mohammad Afzal;Sankalp Gambhir;Ashutosh Gupta;S. Krishna;Ashutosh Trivedi;Alvaro Velasquez
Mohammad Afzal;Sankalp Gambhir;Ashutosh Gupta;S. Krishna;Ashutosh Trivedi;Alvaro Velasquez
中科院分区:
其他
文献类型:
--
作者:
Mohammad Afzal;Sankalp Gambhir;Ashutosh Gupta;S. Krishna;Ashutosh Trivedi;Alvaro Velasquez

文献摘要

相似文献

近年来强化学习的成功得益于适当奖励函数的表征。然而,在此类奖励不直观、难以定义或定义容易出错的情况下,从专家演示中学习奖励信号是有用的。这是逆强化学习(IRL)的关键。虽然以标量奖励信号的形式引发学习要求已被证明是有效的,但这种表示缺乏可解释性并导致学习不透明。我们的目标是通过提出一种新颖的 IRL 方法来缓解这种情况,该方法以流行的形式逻辑(线性时序逻辑(LTL))的形式从专家策略给出的一组跟踪中引出声明性学习要求。所提出方法的一个关键新颖之处是通过一个单词满足 LTL 公式的定量语义,该单词遵循奥卡姆剃刀原理,激励更简单的解释。给定一个由正迹 $P$ 和负迹 $N$ 组成的样本 $S=(P,N)$,所提出的算法自动搜索公式 $\varphi$,该公式提供了样本的最简单解释(在 LTL 的 $GF$ 片段中)。我们已将这种方法作为开源工具 QuantLearn 来实现,以执行基于逻辑的非马尔可夫 IRL。我们的结果证明了所提出的方法从噪声数据中得出直观的基于 LTL 的奖励信号的可行性。
The successes of reinforcement learning in recent years are underpinned by the characterization of suitable reward functions. However, in settings where such rewards are non-intuitive, difficult to define, or otherwise error-prone in their definition, it is useful to instead learn the reward signal from expert demonstrations. This is the crux of inverse reinforcement learning (IRL). While eliciting learning requirements in the form of scalar reward signals has been shown to effective, such representations lack explainability and lead to opaque learning. We aim to mitigate this situation by presenting a novel IRL method for eliciting declarative learning requirements in the form of a popular formal logic -- Linear Temporal Logic (LTL) -- from a set of traces given by the expert policy. A key novelty of the proposed approach is quantitative semantics of satisfaction of an LTL formula by a word that, following Occam's razor principle, incentivizes simpler explanations. Given a sample $S=(P,N)$ consisting of positive traces $P$ and negative traces $N$, the proposed algorithms automate the search for a formula $\varphi$ which provides the simplest explanation (in the $GF$ fragment of LTL) of the samples. We have implemented this approach as an open-source tool QuantLearn to perform logic-based non-Markovian IRL. Our results demonstrate the feasibility of the proposed approach in eliciting intuitive LTL-based reward signals from noisy data.