Opportunistic Qualitative Planning in Stochastic Systems with Incomplete Preferences over Reachability Objectives

Opportunistic Qualitative Planning in Stochastic Systems with Incomplete Preferences over Reachability Objectives
复制标题

DOI:
10.23919/acc55779.2023.10156127
复制
发表时间:
2022-10
期刊:
2023 American Control Conference (ACC)
影响因子:
--
通讯作者:
A. Kulkarni;Jie Fu
A. Kulkarni;Jie Fu
中科院分区:
其他
文献类型:
--
作者:
A. Kulkarni;Jie Fu

文献摘要

相似文献

当不能同时满足所有约束时,偏好在确定要满足哪些目标/约束方面发挥着关键作用。在本文中,我们研究了如何在建模为 MDP 的随机系统中综合偏好满足计划,给定针对时间扩展目标的(可能不完整的)组合偏好模型。我们首先引入新的语义来解释对随机系统无限发挥的偏好。然后,我们引入了“改进”的新概念,以便能够比较无限游戏的两个前缀。基于此,我们定义了两个解决方案概念,称为安全和积极改进(SPI)和安全和几乎确定改进(SASI),分别以正概率和概率一强制改进。我们构建了一种称为改进 MDP 的模型,其中保证至少一项改进的 SPI 和 SASI 策略的综合,减少到在 MDP 中计算积极且几乎肯定获胜的策略。我们提出了一种算法来综合 SPI 和 SASI 策略,从而引发多个连续改进。我们使用机器人运动规划问题演示了所提出的方法。
Preferences play a key role in determining what goals/constraints to satisfy when not all constraints can be satisfied simultaneously. In this paper, we study how to synthesize preference-satisfying plans in a stochastic system modeled as an MDP, given a (possibly incomplete) combinative preference model over temporally extended goals. We start by introducing new semantics to interpret preferences over infinite plays of the stochastic system. Then, we introduce a new notion of ‘improvement’ to enable comparison between two prefixes of an infinite play. Based on this, we define two solution concepts called Safe and Positively Improving (SPI) and Safe and Almost-Sure Improving (SASI) that enforce improvements with a positive probability and with probability one, respectively. We construct a model called an improvement MDP, in which the synthesis of SPI and SASI strategies that guarantee at least one improvement, reduces to computing positive and almost-sure winning strategies in an MDP. We present an algorithm to synthesize the SPI and SASI strategies that induce multiple sequential improvements. We demonstrate the proposed approach using a robot motion planning problem.