Optimal Control of Partially Observable Markov Decision Processes with Finite Linear Temporal Logic Constraints

Optimal Control of Partially Observable Markov Decision Processes with Finite Linear Temporal Logic Constraints
复制标题

DOI:
10.48550/arxiv.2203.09038
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
K. C. Kalagarla;D. Kartik;Dongming Shen;Rahul Jain;A. Nayyar;P. Nuzzo
K. C. Kalagarla;D. Kartik;Dongming Shen;Rahul Jain;A. Nayyar;P. Nuzzo
中科院分区:
其他
文献类型:
--
作者:
K. C. Kalagarla;D. Kartik;Dongming Shen;Rahul Jain;A. Nayyar;P. Nuzzo

文献摘要

相似文献

自主代理通常在状态被部分观察到的情况下运行。除了最大化他们的累积奖励外,智能体还必须执行具有丰富时间和逻辑结构的复杂任务。这些任务可以使用时间逻辑语言,如有限线性时间逻辑(LTL_f)来表示。本文首次提供了一个结构化的框架,用于设计在保证满足时间逻辑规范的概率足够高的情况下使奖励最大化的代理策略。我们将这个问题重新表述为一个受限的部分可观察马尔可夫决策过程(POMDP),并提供了一种新的方法,可以利用现成的无约束的POMDP求解器来求解它。我们的方法保证了高概率的近似最优性和约束满足。我们通过在几个感兴趣的模型上实现它来证明它的有效性。
Autonomous agents often operate in scenarios where the state is partially observed. In addition to maximizing their cumulative reward, agents must execute complex tasks with rich temporal and logical structures. These tasks can be expressed using temporal logic languages like finite linear temporal logic (LTL_f). This paper, for the first time, provides a structured framework for designing agent policies that maximize the reward while ensuring that the probability of satisfying the temporal logic specification is sufficiently high. We reformulate the problem as a constrained partially observable Markov decision process (POMDP) and provide a novel approach that can leverage off-the-shelf unconstrained POMDP solvers for solving it. Our approach guarantees approximate optimality and constraint satisfaction with high probability. We demonstrate its effectiveness by implementing it on several models of interest.