Joint optimisation convex-negative matrix factorisation for multi-modal image collection summarisation based on images and tags

Joint optimisation convex-negative matrix factorisation for multi-modal image collection summarisation based on images and tags
复制标题

基于图像和标签的多模态图像采集摘要联合优化凸负矩阵分解

DOI:
10.1049/iet-cvi.2017.0568
复制
发表时间:
2019
期刊:
IET Comput. Vis.
影响因子:
--
通讯作者:
Hongqi Wang
Hongqi Wang
中科院分区:
--
文献类型:
--
作者:
Wenkai Zhang;Kun Fu;Xian Sun;Yuhang Zhang;Hao Sun;Hongqi Wang

文献摘要

被引文献

相似文献

图像集合摘要旨在表示具有图像和标签的小子集的大规模多模态集合,帮助导航大型图像数据集。大多数现存的方法利用文本到视觉摘要的贡献,忽略了文本主题的视觉贡献。当标签被弱标记时,文本主题不能准确地反映视觉摘要。为了解决这个问题,作者提出了一种新的模型,凸非负矩阵分解的联合优化,它以一种有益的方式结合了图像和标签。目标函数包含视觉错误函数和文本错误函数,共享同一个指示矩阵,连接不同的模态关系。然后,他们提出了一种迭代算法来优化所提出的模型。最后,他们探索了不同视觉特征表示(例如词袋和深度学习)对多模态集合摘要的影响。我们提出的方法,然后使用两个多模态数据集(即MIRFlickr和NUS-WIDE-SCENE)与最先进的算法进行比较。实验结果证明了该方法的有效性。
Image collection summarisation aims to represent a large-scale multi-modal collection with a small subset of images and tags, helping navigate a large image dataset. Most extant methods leverage the contributions of text-to-visual summaries, ignoring the visual contribution to the textual topic. When the tags are weakly labelled, the textual topic cannot accurately reflect the visual summary. To solve this, the authors propose a novel model, joint optimisation of convex non-negative matrix factorisation, which incorporates images and tags in a beneficial way. The objective function contains visual and textual error functions, sharing the same indicator matrix, connecting different modal relations. Then, they propose an iterative algorithm to optimise the proposed model. Finally, they explore the effects of different visual feature representations (e.g. bag-of-words and deep learning) on multi-modal collection summary. Our proposed method is then compared with state-of-the-art algorithms using two multi-modal datasets (i.e. MIRFlickr and NUS-WIDE-SCENE). Experimental results demonstrate the effectiveness of their proposed approach.