Here’s What I’ve Learned: Asking Questions that Reveal Reward Learning

Here’s What I’ve Learned: Asking Questions that Reveal Reward Learning
复制标题

这是我学到的:提出能够奖励学习的问题

DOI:
--
复制
发表时间:
2021
影响因子:
5.1
通讯作者:
Dylan P. Losey
Dylan P. Losey
中科院分区:
--
文献类型:
--
作者:
Soheil Habibian;Ananth Jonnavittula;Dylan P. Losey

文献摘要

参考文献

被引文献

相似文献

机器人可以通过提问向人类学习。在这些问题中,机器人展示了几种不同的行为,并向人类询问他们最喜欢的行为。但是机器人应该如何选择问哪些问题呢?今天的机器人对信息问题进行优化,尽可能有效地主动探测人类的偏好。但是,虽然从机器人的角度来看,信息性问题是有意义的,但人类旁观者可能会发现它们是武断和误导的。例如,考虑一个辅助机器人学习收拾盘子。根据你对之前问题的回答,这个机器人知道应该把每个盘子放在哪里;然而,机器人不确定携带这些盘子的合适高度。一个只针对信息性问题进行优化的机器人只关注这个高度:它会显示盘子在离桌子近或远的地方移动的轨迹,而不管它们是否正确地堆叠盘子。因此,当我们看到这个问题时,我们会错误地认为机器人还在困惑该把盘子堆在哪里!在本文中,我们从人类的角度形式化主动的基于偏好的学习。我们假设——从人类的角度来看——机器人的问题揭示了机器人学到了什么和没有学到什么。我们的洞察力使机器人能够使用问题使他们的学习过程对人类操作员透明。我们开发并测试了一个模型,机器人可以利用这个模型将它们提出的问题与这些问题所揭示的信息联系起来。然后,我们介绍了考虑人类和机器人视角的信息性和揭示性问题之间的权衡:针对这种权衡进行优化的机器人积极地从人类那里收集信息,同时使人类跟上它所学到的知识。我们通过模拟、在线调查和面对面的用户研究来评估我们的方法。我们发现,考虑人类观点的机器人学习速度与最先进的基线一样快,同时还能将它们学到的东西传达给人类操作员。我们用户研究的视频和结果可以在这里找到:https://youtu.be/tC6y_jHN7Vw。
Robots can learn from humans by asking questions. In these questions, the robot demonstrates a few different behaviors and asks the human for their favorite. But how should robots choose which questions to ask? Today’s robots optimize for informative questions that actively probe the human’s preferences as efficiently as possible. But while informative questions make sense from the robot’s perspective, human onlookers may find them arbitrary and misleading. For example, consider an assistive robot learning to put away the dishes. Based on your answers to previous questions this robot knows where it should stack each dish; however, the robot is unsure about right height to carry these dishes. A robot optimizing only for informative questions focuses purely on this height: it shows trajectories that carry the plates near or far from the table, regardless of whether or not they stack the dishes correctly. As a result, when we see this question, we mistakenly think that the robot is still confused about where to stack the dishes! In this article, we formalize active preference-based learning from the human’s perspective. We hypothesize that—from the human’s point-of-view —the robot’s questions reveal what the robot has and has not learned. Our insight enables robots to use questions to make their learning process transparent to the human operator. We develop and test a model that robots can leverage to relate the questions they ask to the information these questions reveal. We then introduce a tradeoff between informative and revealing questions that considers both human and robot perspectives: a robot that optimizes for this tradeoff actively gathers information from the human while simultaneously keeping the human up to date with what it has learned. We evaluate our approach across simulations, online surveys, and in-person user studies. We find that robots, which consider the human’s point of view learn just as quickly as state-of-the-art baselines while also communicating what they have learned to the human operator. Videos of our user studies and results are available here: https://youtu.be/tC6y_jHN7Vw.
DOI: --
发表时间: 2019-04
期刊: --
影响因子: --
作者:
Daniel S. Brown;Wonjoon Goo;P. Nagarajan;S. Niekum
通讯作者: Daniel S. Brown;Wonjoon Goo;P. Nagarajan;S. Niekum
从不同来源的人类反馈中学习奖励函数:最佳地整合演示和偏好
DOI: 10.1177/02783649211041652
发表时间: 2021
期刊: The International Journal of Robotics Research
影响因子: --
作者:
Bıyık, Erdem;Losey, Dylan P.;Palan, Malayandi;Landolfi, Nicholas C.;Shevchuk, Gleb;Sadigh, Dorsa
通讯作者: Sadigh, Dorsa
提出简单的问题:一种用户友好的主动奖励学习方法
DOI: --
发表时间: 2019
期刊: Proceedings of the 3rd Conference on Robot Learning
影响因子: --
作者:
Erdem Biyik, Malayandi Palan
通讯作者: Erdem Biyik, Malayandi Palan
DOI: --
发表时间: 2017-10
期刊: --
影响因子: --
作者:
Jesse Thomason;Aishwarya Padmakumar;Jivko Sinapov;Justin W. Hart;P. Stone;R. Mooney
通讯作者: Jesse Thomason;Aishwarya Padmakumar;Jivko Sinapov;Justin W. Hart;P. Stone;R. Mooney
DOI: 10.1007/s10514-018-9771-0
发表时间: 2019-02-01
期刊: AUTONOMOUS ROBOTS
影响因子: 3.5
作者:
Huang, Sandy H.;Held, David;Dragan, Anca D.
通讯作者: Dragan, Anca D.