Learning Human Objectives by Evaluating Hypothetical Behavior

Learning Human Objectives by Evaluating Hypothetical Behavior
复制标题

DOI:
--
复制
发表时间:
2019-12
期刊:
--
影响因子:
--
通讯作者:
S. Reddy;A. Dragan;S. Levine;S. Legg;J. Leike
S. Reddy;A. Dragan;S. Levine;S. Legg;J. Leike
中科院分区:
其他
文献类型:
--
作者:
S. Reddy;A. Dragan;S. Levine;S. Legg;J. Leike

文献摘要

相似文献

我们试图在具有未知动态、未知奖励函数和未知不安全状态的强化学习设置中使代理行为与用户目标保持一致。用户知道奖励和不安全状态,但是查询用户是昂贵的。为了解决这一挑战,我们提出了一种安全且交互式地学习用户奖励函数模型的算法。我们从初始状态的生成模型和非策略数据训练的前向动态模型开始。我们的方法使用这些模型来合成假设行为,要求用户用奖励标记这些行为,并训练神经网络来预测奖励。关键思想是在不与环境交互的情况下,通过最大化信息价值的可处理代理,从零开始积极地合成假设行为。我们称这种方法为基于轨迹优化的奖励查询综合(ReQueST)。我们在基于状态的2D导航任务和基于图像的赛车视频游戏中模拟用户评估ReQueST。结果表明,ReQueST在学习具有不同初始状态分布的新环境的奖励模型方面明显优于先前的方法。此外,ReQueST安全地训练奖励模型来检测不安全状态,并在部署代理之前纠正奖励黑客行为。
We seek to align agent behavior with a user's objectives in a reinforcement learning setting with unknown dynamics, an unknown reward function, and unknown unsafe states. The user knows the rewards and unsafe states, but querying the user is expensive. To address this challenge, we propose an algorithm that safely and interactively learns a model of the user's reward function. We start with a generative model of initial states and a forward dynamics model trained on off-policy data. Our method uses these models to synthesize hypothetical behaviors, asks the user to label the behaviors with rewards, and trains a neural network to predict the rewards. The key idea is to actively synthesize the hypothetical behaviors from scratch by maximizing tractable proxies for the value of information, without interacting with the environment. We call this method reward query synthesis via trajectory optimization (ReQueST). We evaluate ReQueST with simulated users on a state-based 2D navigation task and the image-based Car Racing video game. The results show that ReQueST significantly outperforms prior methods in learning reward models that transfer to new environments with different initial state distributions. Moreover, ReQueST safely trains the reward model to detect unsafe states, and corrects reward hacking before deploying the agent.