Restraining Bolts for Reinforcement Learning Agents

Restraining Bolts for Reinforcement Learning Agents
复制标题

强化学习代理的约束螺栓

DOI:
10.1609/aaai.v34i09.7114
复制
发表时间:
2020
期刊:
The Knowledge Engineering Review
影响因子:
--
通讯作者:
F. Patrizi
F. Patrizi
中科院分区:
--
文献类型:
--
作者:
Giuseppe De Giacomo;L. Iocchi;Marco Favorito;F. Patrizi

文献摘要

参考文献

被引文献

相似文献

在这项工作中,我们研究了“约束螺栓”的概念,灵感来自科幻小说。我们从世界中提取了两组不同的特征,一组是由代理提取的,另一组是由权威机构对代理的行为施加一些约束规范(“约束螺栓”)提取的。这两组特征以及由此获得的世界模型显然是不相关的,因为它们对独立的各方都有兴趣。然而,它们都反映了同一个世界的方方面面。我们考虑了这样一种情况,在这种情况下,智能体是一组低级(子符号)特征上的强化学习智能体,而约束螺栓是在一组高级符号特征上的有限轨迹f/f上使用线性时间逻辑指定的。我们正式地展示并举例说明,在一般情况下,智能体可以在塑造其目标以适当地(尽可能地)符合约束螺栓规范的同时学习
In this work we have investigated the concept of “restraining bolt”, inspired by Science Fiction. We have two distinct sets of features extracted from the world, one by the agent and one by the authority imposing some restraining specifications on the behaviour of the agent (the “restraining bolt”). The two sets of features and, hence the model of the world attainable from them, are apparently unrelated since of interest to independent parties. However they both account for (aspects of) the same world. We have considered the case in which the agent is a reinforcement learning agent on a set of low-level (subsymbolic) features, while the restraining bolt is specified logically using linear time logic on finite traces f/f over a set of high-level symbolic features. We show formally, and illustrate with examples, that, under general circumstances, the agent can learn while shaping its goals to suitably conform (as much as possible) to the restraining bolt specifications.1
通过 GLTL 实现与环境无关的任务规范
DOI: --
发表时间: 2017
期刊: arXiv.org
影响因子: --
作者:
Littman, Michael L.;Topcu, Ufuk;Fu, Jie;Isbell, Charles;Wen, Min;MacGlashan, James
通讯作者: MacGlashan, James