Model-based motion planning in POMDPs with temporal logic specifications

Model-based motion planning in POMDPs with temporal logic specifications
复制标题

具有时间逻辑规范的 POMDP 中基于模型的运动规划

DOI:
10.1080/01691864.2023.2226191
复制
发表时间:
2023
期刊:
影响因子:
2
通讯作者:
Xiao, Shaoping
Xiao, Shaoping
中科院分区:
计算机科学4区
文献类型:
--
作者:
Li, Junchao;Cai, Mingyu;Wang, Zhaoan;Xiao, Shaoping

文献摘要

参考文献

被引文献

相似文献

部分可观察马尔可夫决策过程(POMDP)已被用作不确定和不完整信息下顺序决策的数学模型。由于 POMDP 中的状态空间是部分可观察的,因此智能体必须根据过去的行动经验和观察的综合信息做出决策。本研究旨在解决概率运动规划问题,其中代理在部分可观察的环境下被分配复杂的任务。我们采用线性时序逻辑(LTL)来制定复杂的任务,然后将其转换为极限确定性广义布奇自动机(LDGBA)。我们将问题重新表述为基于模型检查技术寻找 POMDP 和 LDGBA 乘积的最优策略。本文采用并修改了两种强化学习(RL)方法:值迭代和深度 Q 学习。两者都是基于模型的,因为最优策略是需要更新转移和观察概率的信念状态的函数。我们通过解决两个模拟来说明所提出的方法的适用性,包括不同大小的网格世界问题和 TurtleBot 办公室路径规划问题。
Partially observable Markov decision processes (POMDPs) have been used as mathematical models for sequential decision-making under uncertain and incomplete information. Since the state space is partially observable in a POMDP, the agent has to make a decision based on the integrated information over the past experiences of actions and observations. This study aims to solve probabilistic motion planning problems in which the agent is assigned a complex task under a partially observable environment. We employ linear temporal logic (LTL) to formulate the complex task and then convert it to a limit-deterministic generalized Büchi automaton (LDGBA). We reformulate the problem as finding an optimal policy on the product of POMDP and LDGBA based on model-checking techniques. This paper adopts and modifies two reinforcement learning (RL) approaches: value iteration and deep Q-learning. Both are model-based because the optimal policy is a function of belief states that need transition and observation probabilities to be updated. We illustrate the applicability of the proposed methods by addressing two simulations, including a grid-world problem with various sizes and a TurtleBot office path planning problem.
具有时态逻辑约束的强化学习,用于部分可观察的马尔可夫决策过程
DOI: --
发表时间: 2021
期刊: arXiv.org
影响因子: --
作者:
Yu Wang;A. Bozkurt;Miroslav Pajic
通讯作者: Miroslav Pajic
将 POMDP 抽象与重新规划相结合来解决复杂的、位置相关的传感任务
DOI: --
发表时间: 2013
期刊: AAAI Fall Symposia
影响因子: --
作者:
D. Grady;Mark Moll;L. Kavraki
通讯作者: L. Kavraki
部分可观测马尔可夫决策过程中基于点的模型检验方法
DOI: --
发表时间: 2020
期刊: AAAI Conference on Artificial Intelligence
影响因子: --
作者:
Maxime Bouton;Jana Tumova;Mykel J. Kochenderfer
通讯作者: Mykel J. Kochenderfer
针对机器人应用的具有时序逻辑规范的 POMDP 定性分析
DOI: 10.1109/icra.2015.7139019
发表时间: 2014
期刊: 2015 IEEE International Conference on Robotics and Automation (ICRA)
影响因子: --
作者:
K. Chatterjee;Martin Chmelík;Raghav Gupta;Ayush Kanodia
通讯作者: Ayush Kanodia
用于系统构建和分析的工具和算法
DOI: 10.1007/978-3-642-28756-5_47
发表时间: 2012
期刊: --
影响因子: --
作者:
Basler G
通讯作者: Basler G