Reinforcement Learning with Temporal-Logic-Based Causal Diagrams

Reinforcement Learning with Temporal-Logic-Based Causal Diagrams
复制标题

DOI:
10.48550/arxiv.2306.13732
复制
发表时间:
2023-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Yashi Paliwal;Rajarshi Roy;Jean-Raphael Gaglione;Nasim Baharisangari;D. Neider;Xiaoming Duan;U. Topcu;Zhe Xu
Yashi Paliwal;Rajarshi Roy;Jean-Raphael Gaglione;Nasim Baharisangari;D. Neider;Xiaoming Duan;U. Topcu;Zhe Xu
中科院分区:
其他
文献类型:
--
作者:
Yashi Paliwal;Rajarshi Roy;Jean-Raphael Gaglione;Nasim Baharisangari;D. Neider;Xiaoming Duan;U. Topcu;Zhe Xu

文献摘要

相似文献

我们研究一类强化学习(RL)任务,其中代理的目标是完成临时扩展的目标。在这种情况下,一种常见的方法是将任务表示为确定性有限自动机 (DFA),并将它们集成到 RL 算法的状态空间中。然而,虽然这些机器对奖励函数进行建模,但它们经常忽略有关环境的因果知识。为了解决这个限制,我们在强化学习中提出了基于时间逻辑的因果图(TL-CD),它捕获了环境不同属性之间的时间因果关系。我们利用 TL-CD 设计了一种 RL 算法,其中代理需要显着减少对环境的探索。为此,基于 TL-CD 和任务 DFA,我们确定了代理可以在探索过程中尽早确定预期奖励的配置。通过一系列案例研究,我们展示了使用 TL-CD 的好处,特别是由于减少了对环境的探索,算法可以更快地收敛到最优策略。
We study a class of reinforcement learning (RL) tasks where the objective of the agent is to accomplish temporally extended goals. In this setting, a common approach is to represent the tasks as deterministic finite automata (DFA) and integrate them into the state-space for RL algorithms. However, while these machines model the reward function, they often overlook the causal knowledge about the environment. To address this limitation, we propose the Temporal-Logic-based Causal Diagram (TL-CD) in RL, which captures the temporal causal relationships between different properties of the environment. We exploit the TL-CD to devise an RL algorithm in which an agent requires significantly less exploration of the environment. To this end, based on a TL-CD and a task DFA, we identify configurations where the agent can determine the expected rewards early during an exploration. Through a series of case studies, we demonstrate the benefits of using TL-CDs, particularly the faster convergence of the algorithm to an optimal policy due to reduced exploration of the environment.