Reward Design with Language Models

Reward Design with Language Models
复制标题

DOI:
10.48550/arxiv.2303.00001
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Minae Kwon;Sang Michael Xie;Kalesha Bullard;Dorsa Sadigh
Minae Kwon;Sang Michael Xie;Kalesha Bullard;Dorsa Sadigh
中科院分区:
其他
文献类型:
--
作者:
Minae Kwon;Sang Michael Xie;Kalesha Bullard;Dorsa Sadigh

文献摘要

被引文献

相似文献

强化学习(RL)中的奖励设计具有挑战性,因为通过奖励函数指定人类期望行为的概念可能很困难,或者需要许多专家演示。我们是否可以使用自然语言界面来设计低成本的奖励?本文探讨了如何通过提示大型语言模型(LLM)(如GPT-3)作为代理奖励函数来简化奖励设计,其中用户提供包含期望行为的几个示例(few-shot)或描述(zero-shot)的文本提示。我们的方法利用了RL框架中的代理奖励功能。具体来说,用户在培训开始时指定一次提示。在训练过程中,LLM根据提示描述的期望行为评估RL代理的行为,并输出相应的奖励信号。RL代理然后使用这个奖励来更新它的行为。我们评估了我们的方法是否可以在最后通牒游戏、矩阵游戏和DealOrNoDeal谈判任务中训练与用户目标一致的代理。在这三个任务中,我们表明,用我们的框架训练的强化学习代理与用户的目标很好地一致,并且优于通过监督学习学习的奖励函数训练的强化学习代理
Reward design in reinforcement learning (RL) is challenging since specifying human notions of desired behavior may be difficult via reward functions or require many expert demonstrations. Can we instead cheaply design rewards using a natural language interface? This paper explores how to simplify reward design by prompting a large language model (LLM) such as GPT-3 as a proxy reward function, where the user provides a textual prompt containing a few examples (few-shot) or a description (zero-shot) of the desired behavior. Our approach leverages this proxy reward function in an RL framework. Specifically, users specify a prompt once at the beginning of training. During training, the LLM evaluates an RL agent's behavior against the desired behavior described by the prompt and outputs a corresponding reward signal. The RL agent then uses this reward to update its behavior. We evaluate whether our approach can train agents aligned with user objectives in the Ultimatum Game, matrix games, and the DealOrNoDeal negotiation task. In all three tasks, we show that RL agents trained with our framework are well-aligned with the user's objectives and outperform RL agents trained with reward functions learned via supervised learning