The Assistive Multi-Armed Bandit

The Assistive Multi-Armed Bandit
复制标题

DOI:
10.1109/hri.2019.8673234
复制
发表时间:
2019-01
期刊:
2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI)
影响因子:
--
通讯作者:
Lawrence Chan;Dylan Hadfield-Menell;S. Srinivasa;A. Dragan
Lawrence Chan;Dylan Hadfield-Menell;S. Srinivasa;A. Dragan
中科院分区:
其他
文献类型:
--
作者:
Lawrence Chan;Dylan Hadfield-Menell;S. Srinivasa;A. Dragan

文献摘要

被引文献

相似文献

人类选择中隐含的学习偏好是经济学和计算机科学中研究得很好的问题。然而,大多数工作都假设人类的行为(噪音)与他们的偏好有关。当人们自己了解他们想要什么时,这种方法可能会失败。在这项工作中,我们介绍了辅助多臂强盗,其中一个机器人协助人类玩强盗任务,以最大限度地提高累积奖励。在这个问题中,人类不知道奖励函数,但可以通过从手臂拉动中获得的奖励来学习它;机器人只观察人类拉动哪些手臂,但不观察与每次拉动相关的奖励。我们提供了充分和必要的条件,在这个框架内成功地帮助人类。令人惊讶的是,在孤立的情况下,更好的人类表现并不一定会在机器人的帮助下带来更好的表现:人类策略可以通过有效地将其观察到的奖励传达给机器人来做得更好。我们进行概念验证实验来支持这些结果。我们认为这项工作有助于人类与机器人交互算法背后的理论。
Learning preferences implicit in the choices humans make is a well studied problem in both economics and computer science. However, most work makes the assumption that humans are acting (noisily) optimally with respect to their preferences. Such approaches can fail when people are themselves learning about what they want. In this work, we introduce the assistive multi-armed bandit, where a robot assists a human playing a bandit task to maximize cumulative reward. In this problem, the human does not know the reward function but can learn it through the rewards received from arm pulls; the robot only observes which arms the human pulls but not the reward associated with each pull. We offer sufficient and necessary conditions for successfully assisting the human in this framework. Surprisingly, better human performance in isolation does not necessarily lead to better performance when assisted by the robot: a human policy can do better by effectively communicating its observed rewards to the robot. We conduct proof-of-concept experiments that support these results. We see this work as contributing towards a theory behind algorithms for human-robot interaction.