Learning behaviors via human-delivered discrete feedback: modeling implicit feedback strategies to speed up learning

Learning behaviors via human-delivered discrete feedback: modeling implicit feedback strategies to speed up learning
复制标题

DOI:
10.1007/s10458-015-9283-7
复制
发表时间:
2015
影响因子:
1.9
通讯作者:
R. Loftin;Bei Peng;J. MacGlashan;M. Littman;Matthew E. Taylor;Jeff Huang;D. Roberts
R. Loftin;Bei Peng;J. MacGlashan;M. Littman;Matthew E. Taylor;Jeff Huang;D. Roberts
中科院分区:
计算机科学4区
文献类型:
--
作者:
R. Loftin;Bei Peng;J. MacGlashan;M. Littman;Matthew E. Taylor;Jeff Huang;D. Roberts

文献摘要

被引文献

相似文献

对于真实世界的应用程序,虚拟代理必须能够从非技术用户那里学习新的行为。正反馈和负反馈是训练新行为的直观方式,现有的工作已经提出了从这种反馈中学习的算法。然而,这项工作将反馈视为要最大化的数字奖励,并假设所有培训师都以相同的方式提供反馈。在这项工作中,我们展示了用户可以以许多不同的方式提供反馈,我们将其描述为“培训策略”。具体地,用户可能并不总是响应于动作而给出明确的反馈,并且可能更有可能提供明确的奖励而不是明确的惩罚,反之亦然,使得缺乏反馈本身传达了关于行为的信息。我们提出了一个概率模型的教练反馈,描述了如何教练选择提供明确的奖励和/或明确的惩罚,并在此模型的基础上,开发两种新的学习算法(SABL和I-SABL),考虑到教练的策略,因此可以学习的情况下,没有提供反馈。通过在线用户研究,我们证明,这些算法可以学习更少的反馈比算法的基础上的反馈的数值解释。此外,我们进行了实证分析的培训策略的用户,并影响他们选择的策略的因素。
For real-world applications, virtual agents must be able to learn new behaviors from non-technical users. Positive and negative feedback are an intuitive way to train new behaviors, and existing work has presented algorithms for learning from such feedback. That work, however, treats feedback as numeric reward to be maximized, and assumes that all trainers provide feedback in the same way. In this work, we show that users can provide feedback in many different ways, which we describe as “training strategies.” Specifically, users may not always give explicit feedback in response to an action, and may be more likely to provide explicit reward than explicit punishment, or vice versa, such that the lack of feedback itself conveys information about the behavior. We present a probabilistic model of trainer feedback that describes how a trainer chooses to provide explicit reward and/or explicit punishment and, based on this model, develop two novel learning algorithms (SABL and I-SABL) which take trainer strategy into account, and can therefore learn from cases where no feedback is provided. Through online user studies we demonstrate that these algorithms can learn with less feedback than algorithms based on a numerical interpretation of feedback. Furthermore, we conduct an empirical analysis of the training strategies employed by users, and of factors that can affect their choice of strategy.