Probabilistic Planning with Partially Ordered Preferences over Temporal Goals

Probabilistic Planning with Partially Ordered Preferences over Temporal Goals
复制标题

DOI:
10.1109/icra48891.2023.10160678
复制
发表时间:
2022-09
期刊:
2023 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Hazhar Rahmani;A. Kulkarni;Jie Fu
Hazhar Rahmani;A. Kulkarni;Jie Fu
中科院分区:
其他
文献类型:
--
作者:
Hazhar Rahmani;A. Kulkarni;Jie Fu

文献摘要

相似文献

在这篇文章中,我们研究了随机系统中的规划问题,该系统被建模为马尔可夫决策过程(MDP),其偏好超过时间扩展的目标。以前关于带有偏好的时间规划的工作假设用户偏好形成一个总顺序,这意味着每一对结果彼此都是可比较的。在这项工作中,我们考虑了这样一种情况,即对可能结果的偏好是部分顺序而不是总顺序。我们首先介绍了确定性有限自动机的一种变体,称为偏好DFA,用于指定用户相对于时间扩展目标的偏好。基于序理论,我们将标签MDP中的偏好DFA转化为对策略的偏好关系。在这种处理中,最优策略在MDP中的有限路径上诱导出弱随机非支配概率分布。所提出的规划算法取决于多目标MDP的构建。我们证明了给定偏好规范的弱随机非支配策略在所构造的多目标MDP中是Pareto最优的,反之亦然。在整篇文章中,我们使用一个运行的例子来演示所提出的偏好规范和解决方法。通过算例详细分析了该算法的有效性,并讨论了未来可能的发展方向。
In this paper, we study planning in stochastic systems, modeled as Markov decision processes (MDPs), with preferences over temporally extended goals. Prior work on temporal planning with preferences assumes that the user preferences form a total order, meaning that every pair of outcomes are comparable with each other. In this work, we consider the case where the preferences over possible outcomes are a partial order rather than a total order. We first introduce a variant of deterministic finite automaton, referred to as a preference DFA, for specifying the user's preferences over temporally extended goals. Based on the order theory, we translate the preference DFA to a preference relation over policies for probabilistic planning in a labeled MDP. In this treatment, a most preferred policy induces a weak-stochastic nondominated probability distribution over the finite paths in the MDP. The proposed planning algorithm hinges on the construction of a multi-objective MDP. We prove that a weak-stochastic nondominated policy given the preference specification is Pareto-optimal in the constructed multi-objective MDP, and vice versa. Throughout the paper, we employ a running example to demonstrate the proposed preference specification and solution approaches. We show the efficacy of our algorithm using the example with detailed analysis, and then discuss possible future directions.