Learning Probably Approximately Complete and Safe Action Models for Stochastic Worlds

Learning Probably Approximately Complete and Safe Action Models for Stochastic Worlds
复制标题

DOI:
10.1609/aaai.v36i9.21215
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Brendan Juba;Roni Stern
Brendan Juba;Roni Stern
中科院分区:
其他
文献类型:
--
作者:
Brendan Juba;Roni Stern

文献摘要

相似文献

我们考虑了在未知随机环境中进行计划模型的学习问题,可以使用概率规划域描述语言(PPDDL)来定义。作为输入,我们为我们提供了一组先前执行的轨迹,主要的挑战是学习一个行动模型,该模型具有与创建这些轨迹的策略相似的目标成就概率。为此,我们介绍了PPDDL的一种变体,其中关于过渡概率存在不确定性,该概率由每个因素的间隔指定,其中包含相应的真实过渡概率。然后,我们提出SAM+,一种算法,该算法学习了这种不精确的PPDDL环境模型。 SAM+具有多项式时间和样本复杂性,并确保有高概率,真正的环境确实会被定义的间隔捕获。我们证明,动作模型SAM+输出的目标成就概率几乎比用于产生训练轨迹的政策一样好或更好。然后,我们展示了如何基于具有相似属性的不精确-PPDDL模型产生PPDDL模型。
We consider the problem of learning action models for planning in unknown stochastic environments that can be defined using the Probabilistic Planning Domain Description Language (PPDDL). As input, we are given a set of previously executed trajectories, and the main challenge is to learn an action model that has a similar goal achievement probability to the policies used to create these trajectories. To this end, we introduce a variant of PPDDL in which there is uncertainty about the transition probabilities, specified by an interval for each factor that contains the respective true transition probabilities. Then, we present SAM+, an algorithm that learns such an imprecise-PPDDL environment model. SAM+ has a polynomial time and sample complexity, and guarantees that with high probability, the true environment is indeed captured by the defined intervals. We prove that the action model SAM+ outputs has a goal achievement probability that is almost as good or better than that of the policies used to produced the training trajectories. Then, we show how to produce a PPDDL model based on this imprecise-PPDDL model that has similar properties.