Guiding Pretraining in Reinforcement Learning with Large Language Models

Guiding Pretraining in Reinforcement Learning with Large Language Models
复制标题

DOI:
10.48550/arxiv.2302.06692
复制
发表时间:
2023-02
期刊:
--
影响因子:
--
通讯作者:
Yuqing Du;Olivia Watkins;Zihan Wang;Cédric Colas;Trevor Darrell;P. Abbeel;Abhishek Gupta;Jacob Andreas
Yuqing Du;Olivia Watkins;Zihan Wang;Cédric Colas;Trevor Darrell;P. Abbeel;Abhishek Gupta;Jacob Andreas
中科院分区:
其他
文献类型:
--
作者:
Yuqing Du;Olivia Watkins;Zihan Wang;Cédric Colas;Trevor Darrell;P. Abbeel;Abhishek Gupta;Jacob Andreas

文献摘要

被引文献

相似文献

强化学习算法通常在缺乏密集的、形状良好的奖励函数的情况下挣扎。内在动机的探索方法通过奖励访问新状态或转换的代理来解决这一限制,但这些方法在大环境中提供的好处有限,其中大多数发现的新奇与下游任务无关。我们描述了一种方法,使用背景知识的文本语料库形状探索。这种方法被称为ELLM(用LLM探索),奖励实现语言模型所建议的目标的代理,该语言模型提示了代理当前状态的描述。通过利用大规模语言模型预训练,ELLM引导代理实现对人类有意义和可行的有用行为,而不需要人类参与。我们在Crafter游戏环境和Housekeep机器人模拟器中评估了ELLM,表明ELLM训练的代理在预训练期间具有更好的常识行为覆盖范围,并且通常匹配或提高一系列下游任务的性能。代码可在https://github.com/yuqingd/ellm上获得。
Reinforcement learning algorithms typically struggle in the absence of a dense, well-shaped reward function. Intrinsically motivated exploration methods address this limitation by rewarding agents for visiting novel states or transitions, but these methods offer limited benefits in large environments where most discovered novelty is irrelevant for downstream tasks. We describe a method that uses background knowledge from text corpora to shape exploration. This method, called ELLM (Exploring with LLMs) rewards an agent for achieving goals suggested by a language model prompted with a description of the agent's current state. By leveraging large-scale language model pretraining, ELLM guides agents toward human-meaningful and plausibly useful behaviors without requiring a human in the loop. We evaluate ELLM in the Crafter game environment and the Housekeep robotic simulator, showing that ELLM-trained agents have better coverage of common-sense behaviors during pretraining and usually match or improve performance on a range of downstream tasks. Code available at https://github.com/yuqingd/ellm.