On the Expressivity of Markov Reward

On the Expressivity of Markov Reward
复制标题

DOI:
--
复制
发表时间:
2021-11
期刊:
ArXiv
影响因子:
--
通讯作者:
David Abel;Will Dabney;A. Harutyunyan;Mark K. Ho;M. Littman;Doina Precup;Satinder Singh
David Abel;Will Dabney;A. Harutyunyan;Mark K. Ho;M. Littman;Doina Precup;Satinder Singh
中科院分区:
其他
文献类型:
--
作者:
David Abel;Will Dabney;A. Harutyunyan;Mark K. Ho;M. Littman;Doina Precup;Satinder Singh

文献摘要

被引文献

相似文献

奖励是加强学习代理的驱动力。本文致力于理解奖励的表现力,作为捕获我们希望代理商执行的任务的一种方式。我们围绕可能需要的“任务”的三个新的抽象概念进行构图:(1)一组可接受的行为,(2)对行为的部分顺序,或(3)对轨迹的部分订购。我们的主要结果证明,尽管奖励可以表达许多此类任务,但每个任务类型都存在无马尔可夫奖励功能可以捕获的实例。然后,我们提供一组多项式时间算法,该算法构建了Markov奖励功能,该功能使代理可以优化这三种类型的每种任务,并正确确定何时不存在此类奖励函数。我们以一项证实并说明我们理论发现的实证研究结论。
Reward is the driving force for reinforcement-learning agents. This paper is dedicated to understanding the expressivity of reward as a way to capture tasks that we would want an agent to perform. We frame this study around three new abstract notions of"task"that might be desirable: (1) a set of acceptable behaviors, (2) a partial ordering over behaviors, or (3) a partial ordering over trajectories. Our main results prove that while reward can express many of these tasks, there exist instances of each task type that no Markov reward function can capture. We then provide a set of polynomial-time algorithms that construct a Markov reward function that allows an agent to optimize tasks of each of these three types, and correctly determine when no such reward function exists. We conclude with an empirical study that corroborates and illustrates our theoretical findings.