mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections
复制标题

DOI:
10.48550/arxiv.2205.12005
复制
发表时间:
2022-05
期刊:
--
影响因子:
--
通讯作者:
Chenliang Li;Haiyang Xu;Junfeng Tian;Wei Wang;Ming Yan;Bin Bi;Jiabo Ye;Hehong Chen;Guohai Xu;Zheng-da Cao;Ji Zhang;Songfang Huang;Feiran Huang;Jingren Zhou;Luo Si
Chenliang Li;Haiyang Xu;Junfeng Tian;Wei Wang;Ming Yan;Bin Bi;Jiabo Ye;Hehong Chen;Guohai Xu;Zheng-da Cao;Ji Zhang;Songfang Huang;Feiran Huang;Jingren Zhou;Luo Si
中科院分区:
其他
文献类型:
--
作者:
Chenliang Li;Haiyang Xu;Junfeng Tian;Wei Wang;Ming Yan;Bin Bi;Jiabo Ye;Hehong Chen;Guohai Xu;Zheng-da Cao;Ji Zhang;Songfang Huang;Feiran Huang;Jingren Zhou;Luo Si

文献摘要

被引文献

相似文献

大规模预训练的基础模型已成为构建人工智能(AI)系统的新兴范式,可以快速适应广泛的下游任务。本文介绍了mPLUG,一个新的视觉语言的基础模型,跨模态的理解和生成。现有的预训练模型在跨模态对齐中效率低下,语言信号被长视觉序列淹没。为了解决这两个问题,mPLUG引入了一种有效且高效的视觉语言架构,具有新颖的跨模态跳过连接。mPLUG在大规模图像-文本对上进行端到端的预训练,具有区分性和生成性目标。它在广泛的视觉语言下游任务上取得了最先进的成果,包括图像字幕,图像-文本检索,视觉基础和视觉问答。mPLUG还在视觉语言和视频语言任务上展示了强大的零拍摄可转移性。代码和预训练模型可在https://github.com/alibaba/AliceMind上获得
Large-scale pre-trained foundation models have been an emerging paradigm for building artificial intelligence (AI) systems, which can be quickly adapted to a wide range of downstream tasks. This paper presents mPLUG, a new vision-language foundation model for both cross-modal understanding and generation. Most existing pre-trained models suffer from inefficiency and linguistic signal overwhelmed by long visual sequences in cross-modal alignment. To address both problems, mPLUG introduces an effective and efficient vision-language architecture with novel cross-modal skip-connections.mPLUG is pre-trained end-to-end on large-scale image-text pairs with both discriminative and generative objectives. It achieves state-of-the-art results on a wide range of vision-language downstream tasks, including image captioning, image-text retrieval, visual grounding and visual question answering. mPLUG also demonstrates strong zero-shot transferability on vision-language and video-language tasks. The code and pre-trained models are available at https://github.com/alibaba/AliceMind