Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions

Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions
复制标题

用于在线决策的 Auto-GPT:基准和附加意见

DOI:
10.48550/arxiv.2306.02224
复制
发表时间:
2023
期刊:
ArXiv
影响因子:
--
通讯作者:
Yunzhong He
Yunzhong He
中科院分区:
--
文献类型:
--
作者:
Hui Yang;Sifu Yue;Yunzhong He

文献摘要

参考文献

被引文献

相似文献

Auto-GPT是一种自主代理,它利用了最近在适应大型语言模型(LLM)用于决策任务方面的进展。虽然人们对Auto-GPT定型代理的兴趣与日俱增,但关于Auto-GPT在解决现实世界决策任务中的有效性和灵活性的问题仍然存在。它在现实世界中参与的能力有限,而且没有基准,这些都是造成这些不确定性的原因。在本文中,我们提出了一个全面的基准研究,自动GPT风格的代理在模拟现实世界场景的决策任务。我们的目标是对这个问题有更深入的了解,并了解基于GPT的代理的适应性。我们比较了流行的LLMS,如GPT-4,GPT-3.5,克劳德和维古纳在Auto-GPT风格的决策任务中的表现。此外,我们还引入了附加意见算法,这是一种简单而有效的方法,它将基于监督/模仿的学习者结合到Auto-GPT方案中。这种方法实现了轻量级监督学习,而不需要对基本LLM进行微调。我们通过仔细的基线比较和消融研究证明,附加意见算法显著提高了在线决策基准测试的性能,包括Webshop和ALFWorld。
Auto-GPT is an autonomous agent that leverages recent advancements in adapting Large Language Models (LLMs) for decision-making tasks. While there has been a growing interest in Auto-GPT stypled agents, questions remain regarding the effectiveness and flexibility of Auto-GPT in solving real-world decision-making tasks. Its limited capability for real-world engagement and the absence of benchmarks contribute to these uncertainties. In this paper, we present a comprehensive benchmark study of Auto-GPT styled agents in decision-making tasks that simulate real-world scenarios. Our aim is to gain deeper insights into this problem and understand the adaptability of GPT-based agents. We compare the performance of popular LLMs such as GPT-4, GPT-3.5, Claude, and Vicuna in Auto-GPT styled decision-making tasks. Furthermore, we introduce the Additional Opinions algorithm, an easy and effective method that incorporates supervised/imitation-based learners into the Auto-GPT scheme. This approach enables lightweight supervised learning without requiring fine-tuning of the foundational LLMs. We demonstrate through careful baseline comparisons and ablation studies that the Additional Opinions algorithm significantly enhances performance in online decision-making benchmarks, including WebShop and ALFWorld.
DOI: --
发表时间: 2022-01
期刊: ArXiv
影响因子: --
作者:
Jason Wei;Xuezhi Wang;Dale Schuurmans;Maarten Bosma;E. Chi;F. Xia;Quoc Le;Denny Zhou
通讯作者: Jason Wei;Xuezhi Wang;Dale Schuurmans;Maarten Bosma;E. Chi;F. Xia;Quoc Le;Denny Zhou