Look Twice as Much as You Say: Scene Graph Contrastive Learning for Self-Supervised Image Caption Generation

Look Twice as Much as You Say: Scene Graph Contrastive Learning for Self-Supervised Image Caption Generation
复制标题

DOI:
10.1145/3511808.3557382
复制
发表时间:
2022-10
期刊:
Proceedings of the 31st ACM International Conference on Information & Knowledge Management
影响因子:
--
通讯作者:
Chunhui Zhang;Chao Huang;Youhuan Li;Xiangliang Zhang;Yanfang Ye;Chuxu Zhang
Chunhui Zhang;Chao Huang;Youhuan Li;Xiangliang Zhang;Yanfang Ye;Chuxu Zhang
中科院分区:
其他
文献类型:
--
作者:
Chunhui Zhang;Chao Huang;Youhuan Li;Xiangliang Zhang;Yanfang Ye;Chuxu Zhang

文献摘要

相似文献

图像通常用于各种信息和知识应用,例如广告和推荐。自动化图像标题生成将显著提高图像的可访问性。然而,这种以图像为输入,文本为输出的跨通道任务,学习起来比较困难。虽然现有方法在图像字幕生成方面取得了良好的性能,但它们依赖于需要足够标记数据的监督学习或需要外部数据集作为语言支点的无监督学习。在本文中,我们提出了SGCL,一种新的场景图对比学习模型的自监督图像字幕生成。SGCL采用预训练和微调流水线。具体而言,我们首先应用场景图生成和目标检测方法编码场景图和图像中的视觉信息作为特征表示。最后,设计了一个基于图注意力网络和递归神经网络的解码器网络来生成序列文本作为字幕。为了在SGCL中实现对比学习,我们将场景图增强设计为图像的对比视图,并通过对比学习有效地训练模型,而无需地面真实标签。此外,我们引入了预训练的单词嵌入和上下文投影器来丰富解码器网络中的文本表示,这有利于模型的预训练。一旦预训练阶段完成,我们将进一步微调具有有限标记数据的图像标题生成任务的模型。在基准数据集上进行的大量实验表明,SGCL的性能优于最先进的模型(监督和无监督)。
Images are commonly used for various information and knowledge applications, such as advertising and recommendation. Automating image caption generation will significantly improve image accessibility. This cross-modal task, which takes image as input and text as output, however, is difficult for learning. Though prior methods achieve good performance for image caption generation, they rely on either supervised learning which requires sufficient labeled data or unsupervised learning which needs external dataset as language pivot. In this paper, we propose SGCL, a novel Scene Graph Contrastive Learning model for self-supervised image caption generation. SGCL adopts the pre-training and fine-tuning pipeline. Specifically, we first apply scene graph generation and objection detection method to encode scene graph and visual information in the image as feature representation. Later, a decoder network based on graph attention network and recurrent neural network is further designed to generate sequential text as caption. To enable contrastive learning in SGCL, we design scene graph augmentations as contrastive views of images and train the model effectively without ground-truth labels through contrastive learning. Additionally, we introduce the pre-trained word embedding and the context projector to enrich the text representation in the decoder network, which benefits model pre-training. Once the pre-training phase is finished, we further fine-tune the model for the image caption generation task with limited labeled data. Extensive experiments on benchmark dataset demonstrate that SGCL outperforms state-of-the-art models (both supervised and unsupervised).