Learning Safe Numeric Action Models

Learning Safe Numeric Action Models
复制标题

DOI:
10.1609/aaai.v37i10.26424
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
Argaman Mordoch;Brendan Juba;Roni Stern
Argaman Mordoch;Brendan Juba;Roni Stern
中科院分区:
其他
文献类型:
--
作者:
Argaman Mordoch;Brendan Juba;Roni Stern

文献摘要

相似文献

已经开发了强大的独立域名计划者来解决各种类型的计划问题。这些计划者通常需要在某些计划域描述语言中给出的代理人的行为模型。然而,获得这样的动作模型是一项艰巨的任务。在关键任务领域中,这项任务更具挑战性,在这些任务领域中,试验的方法是学习如何采取行动的方法。在这样的领域中,用于生成计划的行动模型必须是安全的,从某种意义上说,与之生成的计划必须适用并实现其目标。最近已经探索了针对计划的安全行动模型,以用布尔变量充分描述状态。在这项工作中,我们超出了这一限制,并提出了NSAM算法。 NSAM在观察次数中及时运行,在某些条件下,可以保证返回安全的行动模型。我们分析了其最坏情况的样品复杂性,这可能对某些域很棘手。但是,从经验上讲,NSAM可以快速学习一个可以解决域中大多数问题的安全行动模型。
Powerful domain-independent planners have been developed to solve various types of planning problems. These planners often require a model of the acting agent's actions, given in some planning domain description language. Yet obtaining such an action model is a notoriously hard task. This task is even more challenging in mission-critical domains, where a trial-and-error approach to learning how to act is not an option. In such domains, the action model used to generate plans must be safe, in the sense that plans generated with it must be applicable and achieve their goals. Learning safe action models for planning has been recently explored for domains in which states are sufficiently described with Boolean variables. In this work, we go beyond this limitation and propose the NSAM algorithm. NSAM runs in time that is polynomial in the number of observations and, under certain conditions, is guaranteed to return safe action models. We analyze its worst-case sample complexity, which may be intractable for some domains. Empirically, however, NSAM can quickly learn a safe action model that can solve most problems in the domain.