Active Learning of Reward Dynamics from Hierarchical Queries

Active Learning of Reward Dynamics from Hierarchical Queries
复制标题

DOI:
10.1109/iros40897.2019.8968522
复制
发表时间:
2019-11
期刊:
2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Chandrayee Basu;Erdem Biyik;Zhixun He;M. Singhal;Dorsa Sadigh
Chandrayee Basu;Erdem Biyik;Zhixun He;M. Singhal;Dorsa Sadigh
中科院分区:
其他
文献类型:
--
作者:
Chandrayee Basu;Erdem Biyik;Zhixun He;M. Singhal;Dorsa Sadigh

文献摘要

相似文献

让机器人能够在不同的环境中根据人类的偏好采取行动是一项至关重要的任务,机器人专家和机器学习研究人员对此进行了广泛的研究。为了实现这一目标,人类的偏好通常由机器人优化的奖励函数编码。该奖励函数通常是静态的,因为它不随时间或交互而变化。不幸的是,这种静态奖励函数并不总是能够充分捕捉人类偏好,特别是在非静态环境中:人类偏好会随着环境中其他主体的紧急行为而变化。在这项工作中,我们提出学习奖励动态,可以适应具有多个交互代理的非静态环境。我们将奖励动态定义为奖励函数的元组,每个奖励函数对应一种交互模式,以及管理模式之间转换的模式效用函数。因此,奖励动态不仅编码了不同的人类偏好,还编码了偏好的变化方式。我们的贡献在于我们将基于偏好的学习适应分层方法,该方法不仅旨在学习奖励函数,还旨在学习它们如何基于交互而演变。我们推导出人们如何响应分层查询的概率观察模型。我们的算法利用该模型主动选择分层查询,从而最大化从奖励动态的连续假设空间中删除的数量。我们凭经验证明奖励动态可以准确地匹配人类的偏好。
Enabling robots to act according to human preferences across diverse environments is a crucial task, extensively studied by both roboticists and machine learning researchers. To achieve it, human preferences are often encoded by a reward function which the robot optimizes for. This reward function is generally static in the sense that it does not vary with time or the interactions. Unfortunately, such static reward functions do not always adequately capture human preferences, especially, in non-stationary environments: Human preferences change in response to the emergent behaviors of the other agents in the environment. In this work, we propose learning reward dynamics that can adapt in non-stationary environments with several interacting agents. We define reward dynamics as a tuple of reward functions, one for each mode of interaction, and mode-utility functions governing transitions between the modes. Reward dynamics thereby encodes not only different human preferences but also how the preferences change. Our contribution is in the way we adapt preference-based learning into a hierarchical approach that aims at learning not only reward functions but also how they evolve based on interactions. We derive a probabilistic observation model of how people will respond to the hierarchical queries. Our algorithm leverages this model to actively select hierarchical queries that will maximize the volume removed from a continuous hypothesis space of reward dynamics. We empirically demonstrate reward dynamics can match human preferences accurately.