Deep correlation for matching images and text

Deep correlation for matching images and text
复制标题

DOI:
10.1109/cvpr.2015.7298966
复制
发表时间:
2015-06
期刊:
2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
F. Yan;K. Mikolajczyk
F. Yan;K. Mikolajczyk
中科院分区:
其他
文献类型:
--
作者:
F. Yan;K. Mikolajczyk

文献摘要

被引文献

相似文献

本文研究了用深度典型相关分析(DCCA)学习的联合潜在空间中图像和标题的匹配问题。图像和标题数据由基于视觉和文本的深度神经网络的输出表示。当在DCCA框架中使用时,特征的高维性在内存和速度复杂度方面提出了巨大的挑战。我们通过GPU实现解决了这些问题,并提出了处理过拟合的方法。这使得在流行的标题图像匹配基准上评估DCCA方法成为可能。我们将我们的方法与其他最近提出的技术和目前在三个数据集上的最先进的结果进行比较。
This paper addresses the problem of matching images and captions in a joint latent space learnt with deep canonical correlation analysis (DCCA). The image and caption data are represented by the outputs of the vision and text based deep neural networks. The high dimensionality of the features presents a great challenge in terms of memory and speed complexity when used in DCCA framework. We address these problems by a GPU implementation and propose methods to deal with overfitting. This makes it possible to evaluate DCCA approach on popular caption-image matching benchmarks. We compare our approach to other recently proposed techniques and present state of the art results on three datasets.