Robust Web Image Annotation via Exploring Multi-Facet and Structural Knowledge

Robust Web Image Annotation via Exploring Multi-Facet and Structural Knowledge
复制标题

通过探索多方面和结构知识进行稳健的网络图像注释

DOI:
10.1109/tip.2017.2717185
复制
发表时间:
2017-10
影响因子:
10.6
通讯作者:
Li Xuelong
Li Xuelong
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hu Mengqiu;Yang Yang;Shen Fumin;Zhang Luming;Shen Heng Tao;Li Xuelong

文献摘要

参考文献

被引文献

相似文献

近年来,在互联网和数字技术快速发展的推动下,网络图像呈现爆发式增长。鉴于标签可以反映图像的语义内容,自动图像注释可以进一步简化图像语义索引、检索和其他图像管理任务的过程,已成为多媒体领域最重要的研究方向之一。大多数现有的注释方法严重依赖标记良好的训练数据(收集成本昂贵)和/或视觉特征的单一视图(代表性能力不足)。在本文中,受到特征工程(例如 CNN 特征和尺度不变特征变换特征)和网络上取之不尽的图像数据(与噪声和不完整标签相关)的有希望的进步的启发,我们提出了一种有效且鲁棒的方案,称为鲁棒多视图半监督学习(RMSL),以促进图像注释任务。具体来说,我们利用标记图像和未标记图像来揭示内在的数据结构信息。同时,为了全面描述单个数据,我们利用从图像数据的多个方面(即多个视图或特征)导出的相关和互补信息。我们对不同视图的结果设计了一个强大的成对约束,以实现注释的一致性。此外,我们通过 $\ell _{2,p}$ 损失集成了一个鲁棒的分类器学习组件,它可以在学习过程中提供有效的噪声识别能力。最后,我们设计了一种有效的迭代算法来解决 RMSL 中的优化问题。我们对三个不同的数据集进行了全面的实验,结果表明我们提出的方法对于自动图像注释很有前景。
Driven by the rapid development of Internet and digital technologies, we have witnessed the explosive growth of Web images in recent years. Seeing that labels can reflect the semantic contents of the images, automatic image annotation, which can further facilitate the procedure of image semantic indexing, retrieval, and other image management tasks, has become one of the most crucial research directions in multimedia. Most of the existing annotation methods, heavily rely on well-labeled training data (expensive to collect) and/or single view of visual features (insufficient representative power). In this paper, inspired by the promising advance of feature engineering (e.g., CNN feature and scale-invariant feature transform feature) and inexhaustible image data (associated with noisy and incomplete labels) on the Web, we propose an effective and robust scheme, termed robust multi-view semi-supervised learning (RMSL), for facilitating image annotation task. Specifically, we exploit both labeled images and unlabeled images to uncover the intrinsic data structural information. Meanwhile, to comprehensively describe an individual datum, we take advantage of the correlated and complemental information derived from multiple facets of image data (i.e., multiple views or features). We devise a robust pairwise constraint on outcomes of different views to achieve annotation consistency. Furthermore, we integrate a robust classifier learning component via $\ell _{2,p}$ loss, which can provide effective noise identification power during the learning process. Finally, we devise an efficient iterative algorithm to solve the optimization problem in RMSL. We conduct comprehensive experiments on three different data sets, and the results illustrate that our proposed approach is promising for automatic image annotation.
DOI: 10.1007/s11263-013-0658-4
发表时间: 2014-01-01
影响因子: 19.5
作者:
Gong, Yunchao;Ke, Qifa;Lazebnik, Svetlana
通讯作者: Lazebnik, Svetlana
DOI: 10.1109/tip.2016.2619262
发表时间: 2017
影响因子: 10.6
作者:
Li Liu;Zijia Lin;Ling Shao;Fumin Shen;Guiguang Ding;J. Han
通讯作者: Li Liu;Zijia Lin;Ling Shao;Fumin Shen;Guiguang Ding;J. Han
DOI: 10.1145/1291233.1291380
发表时间: 2007-09
期刊: Proceedings of the 15th ACM international conference on Multimedia
影响因子: --
作者:
J. Liu;Bin Wang-;Mingjing Li;Zhiwei Li;Wei-Ying Ma;Hanqing Lu;Songde Ma
通讯作者: J. Liu;Bin Wang-;Mingjing Li;Zhiwei Li;Wei-Ying Ma;Hanqing Lu;Songde Ma
DOI: 10.1109/cvprw.2009.5204255
发表时间: 2009-06
期刊: 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops
影响因子: --
作者:
Xinghua Sun;Ming-yu Chen;Alexander Hauptmann
通讯作者: Xinghua Sun;Ming-yu Chen;Alexander Hauptmann
DOI: 10.1109/cvpr.2011.5995605
发表时间: 2011-06
期刊: CVPR 2011
影响因子: --
作者:
M. Saberian;Hamed Masnadi-Shirazi;N. Vasconcelos
通讯作者: M. Saberian;Hamed Masnadi-Shirazi;N. Vasconcelos