Show and Tell in the Loop: Cross-Modal Circular Correlation Learning

Show and Tell in the Loop: Cross-Modal Circular Correlation Learning
复制标题

DOI:
10.1109/tmm.2018.2877885
复制
发表时间:
2019-06
影响因子:
7.3
通讯作者:
Yuxin Peng;Jinwei Qi
Yuxin Peng;Jinwei Qi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yuxin Peng;Jinwei Qi

文献摘要

被引文献

相似文献

图像、文本等形式多样的多媒体数据量巨大,但其分布和表现形式不一致。人们已经做了很多工作来打破图像和文本之间的界限,以衡量它们之间的相关性。然而,他们集中在公共子空间的转换或单向生成从一个到另一个单独的,不能充分探讨他们的相互作用。值得注意的是,图像和文本之间的双向生成不仅可以提供互补的提示和相互促进学习跨模态相关性,而且跨模态相关性学习可以反馈,为促进跨模态生成过程提供全面的线索。因此,我们的动机,图像和文本之间的信息传输应被视为一个循环的过程,其目的是充分了解他们的潜在相关性,并进一步实现跨模态生成,以产生在一个统一的框架中的真实图像和文本描述。在本文中,我们提出了跨模态循环相关学习方法,通过一个有效的循环学习训练过程,同时执行跨模态相关学习和生成。首先,我们提出了跨模态循环学习模型来循环执行图像到文本的字幕和文本到图像的合成,并学习共同的表示作为往返的桥梁,它可以实现有效的交互,以充分利用潜在的跨模态相关性。其次,提出了一个统一的双向框架来进行跨模态的相互生成,并在一个高效的循环过程中进行训练,以增强共同表示的生成能力,从而可以循环反馈,进一步促进跨模态的相关学习。综上所述,我们在一个统一的框架内,通过循环学习过程,同时执行跨模态检索,图像到文本的标题,和文本到图像的合成,具有高度的可扩展性和通用性,以实现对跨模态数据的通用认知。我们进行了广泛的实验,不仅评估跨模态检索的相关性能,而且还显示MS-COCO数据集上的图像标题和合成的生成有效性。
Multimedia data with various modalities, such as image and text, are huge in quantity but have inconsistent distribution and representation. Many works have been done to break the boundary between image and text to measure their correlation. However, they focus on either the transformation to common subspace or the unidirectional generation from one to another individually, which cannot fully explore their interactions. It is noted that the bidirectional generation between image and text not only can provide complementary hints and mutually boost to learn cross-modal correlation but also cross-modal correlation learning can feed back to give comprehensive clues for promoting the cross-modal generation process. Therefore, we have the motivation that information transmission between image and text should be treated as a circular process, which aims to fully understand their latent correlation, and further realize cross-modal generation to produce both realistic images and text descriptions in a unified framework. In this paper, we propose the cross-modal circular correlation learning approach to perform both cross-modal correlation learning and generation simultaneously through an efficient circular learning training procedure. First, we propose the cross-modal circular learning model to perform an image-to-text caption and text-to-image synthesis circularly and learn common representation as a round-trip bridge, which can realize efficient interactions to fully exploit latent cross-modal correlations. Second, a unified bidirectional framework is proposed to conduct cross-modal mutual generation and is trained in an efficient circular process to enhance the generative ability of common representation, which can feed back circularly to further promote cross-modal correlation learning. In summary, we simultaneously perform cross-modal retrieval, image-to-text caption, and text-to-image synthesis in a unified framework with the circular learning process, which has high scalability and generality to realize universal cognition on the cross-modal data. We conduct extensive experiments to not only evaluate the correlation performance by cross-modal retrieval but also to show the generation effectiveness of both image caption and synthesis on the MS-COCO dataset.