Safe Learning for Uncertainty-Aware Planning via Interval MDP Abstraction

Safe Learning for Uncertainty-Aware Planning via Interval MDP Abstraction
复制标题

DOI:
10.1109/lcsys.2022.3173993
复制
发表时间:
2022-02
影响因子:
3
通讯作者:
Jesse Jiang;Ye Zhao;S. Coogan
Jesse Jiang;Ye Zhao;S. Coogan
中科院分区:
--
文献类型:
--
作者:
Jesse Jiang;Ye Zhao;S. Coogan

文献摘要

被引文献

相似文献

我们研究了部分已知的随机系统对规划规范定义使用语法共同安全的线性时序逻辑(scLTL)的精炼可满足性界限的问题。我们提出了一种基于抽象的方法,迭代生成高置信区间马尔可夫决策过程(IMDP)抽象的系统从高置信度的界限上的未知组件的动态通过高斯过程回归。特别是,我们开发了一种合成策略,通过寻找路径,避免使用产品IMDP违反规范的状态来采样未知的动态。我们还提供了一个启发式的选择在各种候选路径,以最大限度地提高信息增益。最后,我们提出了一个迭代算法来合成一个令人满意的控制策略的产品IMDP系统。我们证明了我们的工作与移动的机器人导航的案例研究。
We study the problem of refining satisfiability bounds for partially-known stochastic systems against planning specifications defined using syntactically co-safe Linear Temporal Logic (scLTL). We propose an abstraction-based approach that iteratively generates high-confidence Interval Markov Decision Process (IMDP) abstractions of the system from high-confidence bounds on the unknown component of the dynamics obtained via Gaussian process regression. In particular, we develop a synthesis strategy to sample the unknown dynamics by finding paths which avoid specification-violating states using a product IMDP. We further provide a heuristic to choose among various candidate paths to maximize the information gain. Finally, we propose an iterative algorithm to synthesize a satisfying control policy for the product IMDP system. We demonstrate our work with a case study on mobile robot navigation.