An Algorithmic Perspective on Imitation Learning

An Algorithmic Perspective on Imitation Learning
复制标题

DOI:
10.1561/2300000053
复制
发表时间:
2018-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Takayuki Osa;J. Pajarinen;G. Neumann;J. Bagnell;P. Abbeel;Jan Peters
Takayuki Osa;J. Pajarinen;G. Neumann;J. Bagnell;P. Abbeel;Jan Peters
中科院分区:
其他
文献类型:
--
作者:
Takayuki Osa;J. Pajarinen;G. Neumann;J. Bagnell;P. Abbeel;Jan Peters

文献摘要

被引文献

相似文献

随着机器人和其他智能代理从简单的环境和问题转移到更复杂、无结构的环境,手动编程它们的行为变得越来越具有挑战性和成本。通常,教师更容易演示所需的行为,而不是试图手动设计它。这个从演示中学习的过程,以及学习这样做的算法,被称为模仿学习。这项工作提供了一个模仿学习的入门。它涵盖了基本的假设、方法和它们之间的关系;为解决问题而开发的丰富的算法集;以及关于有效工具和实施的建议。我们打算将这份文件服务于两个受众。首先,我们希望让机器学习专家熟悉模仿学习的挑战,特别是在机器人学中出现的挑战,以及它与统计监督学习理论和强化学习等更熟悉的框架之间有趣的理论和实践区别。其次,我们希望让机器人专家和应用人工智能专家更广泛地了解可用于模仿学习的框架和工具。我们特别注意模仿学习方法与结构预测方法之间的密切联系,Daume III等人。[2009]。为了组织这次讨论,我们根据以下驱动算法决策的关键标准对模拟学习技术进行分类:1)策略空间的结构。学习的策略是时间索引轨迹(轨迹学习),从观察到行动的映射(所谓的行为克隆[Bain and Sammut,1996]),还是逆最优控制方法中常见的复杂优化或规划问题的结果(Kalman,1964,Moylan and Anderson,1973)。2)培训和测试期间可用的信息。特别是,学习算法是否与教师所拥有的完整状态有关?学习者是否能够与教师互动并收集更正或更多数据?学习者是否有与之交互的系统的(通常是先验的)模型?学习者是否可以使用教师试图优化的奖励(成本)函数?3)成功的概念。不同的算法方法对由此产生的学习行为提供了不同的保证。这些保证的范围从弱的(例如,测量与代理的决策的不一致)到强的(例如,提供关于学习者关于已知或未知的真实成本函数的性能的保证)。我们通过特别注意区分来组织我们的工作(1):将模仿学习分为直接复制期望行为(有时称为行为克隆)和从演示中学习期望行为的隐藏目标(称为逆最优控制或逆强化学习[Russell,1998])。在后一种情况下,行为是为学习者面临的每个新实例解决的优化问题的结果。除了方法分析,我们还讨论了实践者在选择模仿学习方法时必须做出的设计决定。此外,应用实例--例如会打乒乓球的机器人[Kober and Peters,2009]、下围棋的程序[Silver et al.,2016]以及理解自然语言的系统[Win et al.,2015]--说明了不同形式的模仿学习背后的属性和动机。最后,我们提出了一组开放的问题,并指出了机器学习未来可能的研究方向。
As robots and other intelligent agents move from simple environments and problems to more complex, unstructured settings, manually programming their behavior has become increasingly challenging and expensive. Often, it is easier for a teacher to demonstrate a desired behavior rather than attempt to manually engineer it. This process of learning from demonstrations, and the study of algorithms to do so, is called imitation learning. This work provides an introduction to imitation learning. It covers the underlying assumptions, approaches, and how they relate; the rich set of algorithms developed to tackle the problem; and advice on effective tools and implementation. We intend this paper to serve two audiences. First, we want to familiarize machine learning experts with the challenges of imitation learning, particularly those arising in robotics, and the interesting theoretical and practical distinctions between it and more familiar frameworks like statistical supervised learning theory and reinforcement learning. Second, we want to give roboticists and experts in applied artificial intelligence a broader appreciation for the frameworks and tools available for imitation learning. We pay particular attention to the intimate connection between imitation learning approaches and those of structured prediction Daume III et al. [2009]. To structure this discussion, we categorize imitation learning techniques based on the following key criteria which drive algorithmic decisions: 1) The structure of the policy space. Is the learned policy a time-index trajectory (trajectory learning), a mapping from observations to actions (so called behavioral cloning [Bain and Sammut, 1996]), or the result of a complex optimization or planning problem at each execution as is common in inverse optimal control methods [Kalman, 1964, Moylan and Anderson, 1973]. 2) The information available during training and testing. In particular, is the learning algorithm privy to the full state that the teacher possess? Is the learner able to interact with the teacher and gather corrections or more data? Does the learner have a (typically a priori) model of the system with which it interacts? Does the learner have access to the reward (cost) function that the teacher is attempting to optimize? 3) The notion of success. Different algorithmic approaches provide varying guarantees on the resulting learned behavior. These guarantees range from weaker (e.g., measuring disagreement with the agent’s decision) to stronger (e.g., providing guarantees on the performance of the learner with respect to a true cost function, either known or unknown). We organize our work by paying particular attention to distinction (1): dividing imitation learning into directly replicating desired behavior (sometimes called behavioral cloning) and learning the hidden objectives of the desired behavior from demonstrations (called inverse optimal control or inverse reinforcement learning [Russell, 1998]). In the latter case, behavior arises as the result of an optimization problem solved for each new instance that the learner faces. In addition to method analysis, we discuss the design decisions a practitioner must make when selecting an imitation learning approach. Moreover, application examples—such as robots that play table tennis [Kober and Peters, 2009], programs that play the game of Go [Silver et al., 2016], and systems that understand natural language [Wen et al., 2015]— illustrate the properties and motivations behind different forms of imitation learning. We conclude by presenting a set of open questions and point towards possible future research directions for machine learning.