Eventual Discounting Temporal Logic Counterfactual Experience Replay

Eventual Discounting Temporal Logic Counterfactual Experience Replay
复制标题

DOI:
10.48550/arxiv.2303.02135
复制
发表时间:
2023-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Cameron Voloshin;Abhinav Verma;Yisong Yue
Cameron Voloshin;Abhinav Verma;Yisong Yue
中科院分区:
其他
文献类型:
--
作者:
Cameron Voloshin;Abhinav Verma;Yisong Yue

文献摘要

相似文献

线性时间逻辑(LTL)提供了一种简化的方式来指定策略优化任务,否则可能很难用标量奖励功能来描述。但是,标准的RL框架可能太近视而无法找到最大的LTL满足政策。本文做出了两个贡献。首先,我们使用一种我们称为最终折扣的技术开发了一个新的基于价值功能的代理,根据该技术,人们可以找到满足LTL规范具有最高可实现概率的政策。其次,我们开发了一种新的体验重播方法,用于通过反事实推理以满足LTL规范的不同方式来生成政策推广的违反政策数据。我们在离散和连续的国家行动空间中进行的实验确认了我们的反事实经验重播方法的有效性。
Linear temporal logic (LTL) offers a simplified way of specifying tasks for policy optimization that may otherwise be difficult to describe with scalar reward functions. However, the standard RL framework can be too myopic to find maximally LTL satisfying policies. This paper makes two contributions. First, we develop a new value-function based proxy, using a technique we call eventual discounting, under which one can find policies that satisfy the LTL specification with highest achievable probability. Second, we develop a new experience replay method for generating off-policy data from on-policy rollouts via counterfactual reasoning on different ways of satisfying the LTL specification. Our experiments, conducted in both discrete and continuous state-action spaces, confirm the effectiveness of our counterfactual experience replay approach.