WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
复制标题

DOI:
10.48550/arxiv.2207.01206
复制
发表时间:
2022-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Shunyu Yao;Howard Chen;John Yang;Karthik Narasimhan
Shunyu Yao;Howard Chen;John Yang;Karthik Narasimhan
中科院分区:
其他
文献类型:
--
作者:
Shunyu Yao;Howard Chen;John Yang;Karthik Narasimhan

文献摘要

相似文献

现有的基准接地语言在交互式环境中要么缺乏现实世界的语言元素,或证明难以扩大规模,由于大量的人类参与收集数据或反馈信号。为了弥合这一差距,我们开发了WebShop --一个模拟的电子商务网站环境,拥有118万美元的真实世界产品和12,087美元的众包文本说明。给定指定产品需求的文本指令,代理需要浏览多种类型的网页并发布各种操作来查找、定制和购买物品。WebShop为语言基础提供了几个挑战,包括理解组合指令,查询(重新)公式化,理解和处理网页中的嘈杂文本,以及执行战略探索。我们收集了超过1,600美元的人类演示任务,并使用强化学习,模仿学习和预训练的图像和语言模型来训练和评估各种代理。我们的最佳模型实现了29\%$的任务成功率,优于基于规则的启发式算法(9.6\%$),但远低于人类专家的性能(59\%$)。我们还分析了代理和人类的轨迹,并消融各种模型组件,为开发具有更强语言理解和决策能力的未来代理提供见解。最后,我们表明,在网上商店训练的代理表现出非平凡的模拟到真实的转让时,评估amazon.com和ebay.com,表明网上商店在开发实用的基于Web的代理,可以在野外操作的潜在价值。
Existing benchmarks for grounding language in interactive environments either lack real-world linguistic elements, or prove difficult to scale up due to substantial human involvement in the collection of data or feedback signals. To bridge this gap, we develop WebShop -- a simulated e-commerce website environment with $1.18$ million real-world products and $12,087$ crowd-sourced text instructions. Given a text instruction specifying a product requirement, an agent needs to navigate multiple types of webpages and issue diverse actions to find, customize, and purchase an item. WebShop provides several challenges for language grounding including understanding compositional instructions, query (re-)formulation, comprehending and acting on noisy text in webpages, and performing strategic exploration. We collect over $1,600$ human demonstrations for the task, and train and evaluate a diverse range of agents using reinforcement learning, imitation learning, and pre-trained image and language models. Our best model achieves a task success rate of $29\%$, which outperforms rule-based heuristics ($9.6\%$) but is far lower than human expert performance ($59\%$). We also analyze agent and human trajectories and ablate various model components to provide insights for developing future agents with stronger language understanding and decision making abilities. Finally, we show that agents trained on WebShop exhibit non-trivial sim-to-real transfer when evaluated on amazon.com and ebay.com, indicating the potential value of WebShop in developing practical web-based agents that can operate in the wild.