Cross-modality Bridging and Knowledge Transferring for Image Understanding

Cross-modality Bridging and Knowledge Transferring for Image Understanding
复制标题

用于图像理解的跨模态桥接和知识传输

DOI:
10.1109/tmm.2019.2903448
复制
发表时间:
2019
影响因子:
7.3
通讯作者:
Qionghai Dai
Qionghai Dai
中科院分区:
计算机科学1区
文献类型:
--
作者:
Chenggang Yan;Liang Li;Chunjie Zhang;Bingtao Liu;Yongdong Zhang;Qionghai Dai

文献摘要

被引文献

相似文献

Web图像的理解一直是人工智能和多媒体内容分析领域的研究热点。网络图像是由各种复杂的前景和背景组成的,这使得设计一个准确和鲁棒的学习算法成为一项具有挑战性的任务。为了解决上述重要问题,首先,我们学习了一个跨模态桥接字典,用于深入和完整地理解大量的Web图像。该算法将视觉特征融入到语义概念概率分布中,可以在保持局部几何结构的同时构建图像的全局语义描述。为了发现和建模类别内和类别间的发生模式,引入多任务学习来构造目标公式,并引入Capped-$\ell _{1}$惩罚,该方法能够以更高的概率获得最优解,性能优于传统的基于凸函数的方法.其次,我们提出了一个基于知识的概念转换算法来发现不同类别之间的潜在关系。这种分布概率在类别间的转移可以带来更健壮的全局特征表示,并使图像语义表示随着场景的增大而具有更好的泛化能力。在ImageNet、Caltech-256、SUN 397和Scene 15数据集上与经典方法进行的实验比较和性能讨论表明,我们提出的方法在三个传统图像理解任务中是有效的。
The understanding of web images has been a hot research topic in both artificial intelligence and multimedia content analysis domains. The web images are composed of various complex foregrounds and backgrounds, which makes the design of an accurate and robust learning algorithm a challenging task. To solve the above significant problem, first, we learn a cross-modality bridging dictionary for the deep and complete understanding of a vast quantity of web images. The proposed algorithm leverages the visual features into the semantic concept probability distribution, which can construct a global semantic description for images while preserving the local geometric structure. To discover and model the occurrence patterns between intra- and inter-categories, multi-task learning is introduced for formulating the objective formulation with Capped-$\ell _{1}$ penalty, which can obtain the optimal solution with a higher probability and outperform the traditional convex function-based methods. Second, we propose a knowledge-based concept transferring algorithm to discover the underlying relations of different categories. This distribution probability transferring among categories can bring the more robust global feature representation, and enable the image semantic representation to generalize better as the scenario becomes larger. Experimental comparisons and performance discussion with classical methods on the ImageNet, Caltech-256, SUN397, and Scene15 datasets show the effectiveness of our proposed method at three traditional image understanding tasks.