LESS is More: Rethinking Probabilistic Models of Human Behavior

LESS is More: Rethinking Probabilistic Models of Human Behavior
复制标题

少即是多:重新思考人类行为的概率模型

DOI:
10.1145/3319502.3374811
复制
发表时间:
2020
期刊:
International Conference on Human-Robot Interaction (HRI
影响因子:
--
通讯作者:
Dragan, Anca D.
Dragan, Anca D.
中科院分区:
--
文献类型:
--
作者:
Bobu, Andreea;Scobee, Dexter R.;Fisac, Jaime F.;Sastry, S. Shankar;Dragan, Anca D.

文献摘要

参考文献

被引文献

相似文献

机器人需要人类行为模型来推断人类的目标和偏好,并预测人们会做什么。一个常见的模型是玻尔兹曼噪声-理性决策模型,它假设人们近似地优化奖励函数,并选择与他们的指数奖励成比例的轨迹。虽然这个模型在各种机器人领域都取得了成功,但它的根源在于计量经济学,以及对不同离散选项之间的决策建模,每个选项都有自己的效用或奖励。相比之下,人类轨迹位于连续空间中,具有影响奖励函数的连续值特征。我们建议,现在是时候重新思考玻尔兹曼模型,并从头开始设计它,以在这样的轨迹空间上运行。我们引入了一个模型,明确说明轨迹之间的距离,而不仅仅是他们的奖励。不是每个轨迹独立地影响决策,而是类似的轨迹一起影响决策。我们首先展示了我们的模型在用户研究中更好地解释了人类行为。然后,我们分析了这对机器人推理的影响,首先是在玩具环境中,我们有地面真理,并找到更准确的推理,最后是从用户演示中学习的7自由度机器人手臂。
Robots need models of human behavior for both inferring human goals and preferences, and predicting what people will do. A common model is the Boltzmann noisily-rational decision model, which assumes people approximately optimize a reward function and choose trajectories in proportion to their exponentiated reward. While this model has been successful in a variety of robotics domains, its roots lie in econometrics, and in modeling decisions among different discrete options, each with its own utility or reward. In contrast, human trajectories lie in a continuous space, with continuous-valued features that influence the reward function. We propose that it is time to rethink the Boltzmann model, and design it from the ground up to operate over such trajectory spaces. We introduce a model that explicitly accounts for distances between trajectories, rather than only their rewards. Rather than each trajectory affecting the decision independently, similar trajectories now affect the decision together. We start by showing that our model better explains human behavior in a user study. We then analyze the implications this has for robot inference, first in toy environments where we have ground truth and find more accurate inference, and finally for a 7DOF robot arm learning from user demonstrations.
目标推断作为逆向规划
DOI: --
发表时间: 2007
期刊:
影响因子: --
作者:
Chris L. Baker;J. Tenenbaum;R. Saxe
通讯作者: R. Saxe
DOI: 10.1201/ebk1439808184
发表时间: 2010-06
期刊: --
影响因子: --
作者:
S. Wellek
通讯作者: S. Wellek
DOI: 10.3758/s13428-015-0642-8
发表时间: 2016-09-01
影响因子: 5.4
作者:
Gureckis, Todd M.;Martin, Jay;Chan, Patricia
通讯作者: Chan, Patricia