PaLI: A Jointly-Scaled Multilingual Language-Image Model

PaLI: A Jointly-Scaled Multilingual Language-Image Model
复制标题

DOI:
10.48550/arxiv.2209.06794
复制
发表时间:
2022-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Xi Chen;Xiao Wang;Soravit Changpinyo;A. Piergiovanni;Piotr Padlewski;Daniel M. Salz;Sebastian Goodman;Adam Grycner;Basil Mustafa;Lucas Beyer;Alexander Kolesnikov;J. Puigcerver;Nan Ding;Keran Rong;Hassan Akbari;Gaurav Mishra;Linting Xue;Ashish V. Thapliyal;James Bradbury;Weicheng Kuo;Mojtaba Seyedhosseini;Chao Jia;Burcu Karagol Ayan;C. Riquelme;A. Steiner;A. Angelova;Xiaohua Zhai;N. Houlsby;Radu Soricut
Xi Chen;Xiao Wang;Soravit Changpinyo;A. Piergiovanni;Piotr Padlewski;Daniel M. Salz;Sebastian Goodman;Adam Grycner;Basil Mustafa;Lucas Beyer;Alexander Kolesnikov;J. Puigcerver;Nan Ding;Keran Rong;Hassan Akbari;Gaurav Mishra;Linting Xue;Ashish V. Thapliyal;James Bradbury;Weicheng Kuo;Mojtaba Seyedhosseini;Chao Jia;Burcu Karagol Ayan;C. Riquelme;A. Steiner;A. Angelova;Xiaohua Zhai;N. Houlsby;Radu Soricut
中科院分区:
其他
文献类型:
--
作者:
Xi Chen;Xiao Wang;Soravit Changpinyo;A. Piergiovanni;Piotr Padlewski;Daniel M. Salz;Sebastian Goodman;Adam Grycner;Basil Mustafa;Lucas Beyer;Alexander Kolesnikov;J. Puigcerver;Nan Ding;Keran Rong;Hassan Akbari;Gaurav Mishra;Linting Xue;Ashish V. Thapliyal;James Bradbury;Weicheng Kuo;Mojtaba Seyedhosseini;Chao Jia;Burcu Karagol Ayan;C. Riquelme;A. Steiner;A. Angelova;Xiaohua Zhai;N. Houlsby;Radu Soricut

文献摘要

被引文献

相似文献

有效的扩展和灵活的任务接口使大型语言模型能够在许多任务中表现出色。我们提出了PaLI(路径语言和图像模型),一个将这种方法扩展到语言和视觉的联合建模的模型。PaLI基于视觉和文本输入生成文本,并通过该界面以多种语言执行许多视觉、语言和多模态任务。为了训练PaLI,我们使用了大型预训练的编码器-解码器语言模型和视觉变形器(ViTs)。这使我们能够利用他们现有的能力,并利用培训他们的大量成本。我们发现视觉和语言组件的联合缩放是很重要的。由于现有的语言变形器比它们的视觉对应物要大得多,我们训练了一个大的,40亿个参数的视觉模型(ViT-e)来量化更大容量视觉模型的好处。为了训练巴利语,我们基于一个新的图像-文本训练集创建了一个大型的多语言混合预训练任务,该训练集包含100多种语言的10B图像和文本。PaLI在多种视觉和语言任务(如字幕、视觉问答、场景文本理解)中实现了最先进的技术,同时保留了简单、模块化和可扩展的设计。
Effective scaling and a flexible task interface enable large language models to excel at many tasks. We present PaLI (Pathways Language and Image model), a model that extends this approach to the joint modeling of language and vision. PaLI generates text based on visual and textual inputs, and with this interface performs many vision, language, and multimodal tasks, in many languages. To train PaLI, we make use of large pre-trained encoder-decoder language models and Vision Transformers (ViTs). This allows us to capitalize on their existing capabilities and leverage the substantial cost of training them. We find that joint scaling of the vision and language components is important. Since existing Transformers for language are much larger than their vision counterparts, we train a large, 4-billion parameter ViT (ViT-e) to quantify the benefits from even larger-capacity vision models. To train PaLI, we create a large multilingual mix of pretraining tasks, based on a new image-text training set containing 10B images and texts in over 100 languages. PaLI achieves state-of-the-art in multiple vision and language tasks (such as captioning, visual question-answering, scene-text understanding), while retaining a simple, modular, and scalable design.