PaLM-E: An Embodied Multimodal Language Model

PaLM-E: An Embodied Multimodal Language Model
复制标题

DOI:
10.48550/arxiv.2303.03378
复制
发表时间:
2023-03
期刊:
--
影响因子:
--
通讯作者:
Danny Driess;F. Xia;Mehdi S. M. Sajjadi;Corey Lynch;Aakanksha Chowdhery;Brian Ichter;Ayzaan Wahid;Jonathan Tompson;Q. Vuong;Tianhe Yu;Wenlong Huang;Yevgen Chebotar;P. Sermanet;Daniel Duckworth;S. Levine;Vincent Vanhoucke;Karol Hausman;Marc Toussaint;Klaus Greff;Andy Zeng;Igor Mordatch;Peter R. Florence
Danny Driess;F. Xia;Mehdi S. M. Sajjadi;Corey Lynch;Aakanksha Chowdhery;Brian Ichter;Ayzaan Wahid;Jonathan Tompson;Q. Vuong;Tianhe Yu;Wenlong Huang;Yevgen Chebotar;P. Sermanet;Daniel Duckworth;S. Levine;Vincent Vanhoucke;Karol Hausman;Marc Toussaint;Klaus Greff;Andy Zeng;Igor Mordatch;Peter R. Florence
中科院分区:
其他
文献类型:
--
作者:
Danny Driess;F. Xia;Mehdi S. M. Sajjadi;Corey Lynch;Aakanksha Chowdhery;Brian Ichter;Ayzaan Wahid;Jonathan Tompson;Q. Vuong;Tianhe Yu;Wenlong Huang;Yevgen Chebotar;P. Sermanet;Daniel Duckworth;S. Levine;Vincent Vanhoucke;Karol Hausman;Marc Toussaint;Klaus Greff;Andy Zeng;Igor Mordatch;Peter R. Florence

文献摘要

被引文献

相似文献

大型语言模型擅长各种复杂任务。然而,在真实的世界中实现一般推断,例如,对于机器人技术的问题,提出了接地的挑战。我们提出体现语言模型,直接将现实世界的连续传感器模态到语言模型,从而建立单词和感知之间的联系。我们的体现语言模型的输入是多模态的句子,交错视觉,连续状态估计和文本输入编码。我们结合预先训练的大型语言模型,对这些编码进行端到端的训练,用于多个具体任务,包括顺序机器人操作规划,视觉问答和字幕。我们的评估表明,PaLM-E,一个单一的大型体现多模态模型,可以解决各种体现推理任务,从各种观察模式,在多个实施例,并进一步表现出积极的转移:该模型受益于跨互联网规模的语言,视觉和视觉语言领域的各种联合训练。我们最大的模型PaLM-E-562 B具有562 B参数,除了接受机器人任务的训练外,它还是一个视觉语言通才,在OK-VQA上具有最先进的性能,并随着规模的增加而保留了通才语言能力。
Large language models excel at a wide range of complex tasks. However, enabling general inference in the real world, e.g., for robotics problems, raises the challenge of grounding. We propose embodied language models to directly incorporate real-world continuous sensor modalities into language models and thereby establish the link between words and percepts. Input to our embodied language model are multi-modal sentences that interleave visual, continuous state estimation, and textual input encodings. We train these encodings end-to-end, in conjunction with a pre-trained large language model, for multiple embodied tasks including sequential robotic manipulation planning, visual question answering, and captioning. Our evaluations show that PaLM-E, a single large embodied multimodal model, can address a variety of embodied reasoning tasks, from a variety of observation modalities, on multiple embodiments, and further, exhibits positive transfer: the model benefits from diverse joint training across internet-scale language, vision, and visual-language domains. Our largest model, PaLM-E-562B with 562B parameters, in addition to being trained on robotics tasks, is a visual-language generalist with state-of-the-art performance on OK-VQA, and retains generalist language capabilities with increasing scale.