Separating Skills from Preference: Using Learning to Program by Reward

Separating Skills from Preference: Using Learning to Program by Reward
复制标题

将技能与偏好分开:利用学习通过奖励进行编程

DOI:
--
复制
发表时间:
2002
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
P. Langley
P. Langley
中科院分区:
--
文献类型:
--
作者:
D. Shapiro;P. Langley

文献摘要

被引文献

相似文献

人工智能体的开发人员通常认为,我们只能通过昂贵的实现新技能的过程来指定代理行为。本文提出了一种分离假说:个体之间的行为差异是由于对同一组技能的不同偏好的作用。我们测试这个假设在一个模拟的汽车领域使用强化学习算法来诱导车辆控制策略,给定一个结构化的技能,包含选项,和用户提供的奖励函数。我们表明,定性不同的奖励功能产生代理人在同一组技能的定性不同的行为。这就引出了一个新的开发隐喻,我们称之为奖励式编程。1.动机和背景在许多领域,人类表现出复杂的身体行为,让他们完成复杂的任务。研究人员已经探索了两种主要的方法来学习这些行为,每种方法都与不同类别的表征形式主义有关。一种范例将控制知识编码为规则或类似结构(例如,Laird & Rosenbloom,1990; Sammut,1996),说明了执行动作的条件。一个替代的框架指定了一些将状态-动作对映射到数值实用程序上的函数(例如,Watkins & Dayan,1992),然后用于在动作中进行选择。这两种方法都一再证明了它们在广泛领域中学习有用的控制策略的能力,但每种方法都最自然地适用于智能行为的不同方面。这个想法在游戏中得到了最好的说明,开发人员经常使用规则或其他逻辑约束来指定哪些移动是法律的,但调用数值评估函数来选择它们。我们声称,类似的分工将证明是有用的反应控制政策的研究,包括学习这样的政策,从代理经验。在本文中,我们假设代理已经可以访问一组逻辑规则,限制允许的行动,但它必须学习延迟奖励的剩余选项的值。在其他地方(Shapiro et al.,2001),我们已经表明,这种背景知识的使用可以大大加快学习控制策略的过程。在这里,我们关注的是一个不同的主张:为学习代理提供不同的奖励信号可以导致各种各样的行为,这些行为仍然具有相同的整体结构。这种学习方法--我们称之为奖励编程--应该证明在构建计算机游戏的模拟代理、支持必须在某些约束条件下运行的个性化服务以及许多其他任务方面是有用的。在下面的页面中,我们报告了这个通用框架的一个实例,我们已经在一个名为Icarus的物理代理架构中进行了铸造。我们开始描述架构的逻辑形式主义编码层次技能,从驾驶汽车的任务的例子。然后,我们转向伊卡洛斯用于在适用技能中进行选择的价值函数,以及使用延迟奖励来更新这些函数的算法。在此之后,我们提出了实验研究,旨在验证我们的假设,提供这样一个系统,不同的奖励可以产生独特的,但仍然可行的政策。最后,我们研究了学习复杂技能的其他一些方法,并为这一主题的进一步研究提出了方向。2.伊卡洛斯语言(英语:Icarus Language)是一种用于指定人工智能体学习行为的语言。它的结构是双重动机的愿望,以建立实际的代理应用程序,并希望支持政策学习的计算效率的方式。我们对第一个目标的回应是为伊卡洛斯提供强有力的表现。然而,对快速学习的期望建议了一种更简单的格式,该格式提供了到马尔可夫决策过程(MDP)模型的清晰映射,因为MDP提供了用于开发学习算法的概念框架,
Developers of artificial agents commonly take the view that we can only specify agent behavior via the expensive process of implementing new skills. This paper offers an alternative expressed by the separation hypothesis: that the behavioral differences among individuals are due to the action of distinct preferences over the same set of skills. We test this hypothesis in a simulated automotive domain by using a reinforcement learning algorithm to induce vehicle control policies, given a structured skill for driving that contains options, and a user-supplied reward function. We show that qualitatively distinct reward functions produce agents with qualitatively distinct behavior over the same set of skills. This leads to a new development metaphor we call Ôprogramming by rewardÕ. 1. Motivation and Background In many domains, humans exhibit complex physical behaviors that let them accomplish sophisticated tasks. Researchers have explored two main approaches to learning such behaviors, each associated with a different class of representational formalisms. One paradigm encodes control knowledge as rules or similar structures (e.g., Laird & Rosenbloom, 1990; Sammut, 1996) that state conditions under which to execute actions. An alternative framework instead specifies some function that maps state-action pairs onto a numeric utility (e.g., Watkins & Dayan, 1992), which is then used to select among actions. Both approaches have repeatedly demonstrated their ability to learn useful control policies across a broad range of domains, yet each lends itself most naturally to different aspects of intelligent behavior. This idea is best illustrated by work on game playing, where developers regularly use rules or other logical constraints to specify which moves are legal but invoke numeric evaluation functions to select among them. We claim that a similar division of labor will prove useful in research on policies for reactive control, including learning such policies from agent experience. In this paper, we assume that an agent already has access to a set of logical rules that constrain the allowable actions, but that it must learn the value of its remaining options from delayed reward. Elsewhere (Shapiro et al., 2001), we have shown that this use of background knowledge can greatly speed the process of learning control policies. Here we focus on a different claim: that providing a learning agent with different reward signals can lead to a great variety of behaviors that still share the same overall structure. This approach to learning — which we call programming by reward -should prove useful in constructing simulated agents for computer games, in supporting personalized services that must operate within certain constraints, and many other tasks. In the following pages, we report one instance of this general framework, which we have cast in an architecture for physical agents called Icarus. We begin by describing the architecture’s logical formalism for encoding hierarchical skills, taking examples from the task of driving an automobile. We then turn to the value functions that Icarus uses to select among applicable skills and its algorithm for using delayed rewards to update these functions. After this, we present experimental studies designed to test our hypothesis that providing such a system with different rewards can produce distinctive yet still viable policies. Finally, we examine some other approaches to learning complex skills and suggest directions for additional research on this topic. 2. The Icarus Language Icarus is a language for specifying the behavior of artificial agents that learn. Its structure is dually motivated by the desire to build practical agent applications and the desire to support policy learning in a computationally efficient way. We responded to the first goal by providing Icarus with powerful representations. However, the desire for rapid learning suggests a simpler format that offers a clear mapping into the Markov decision process (MDP) model, since MDPs provide a conceptual framework for developing learning algorithms,