Learning Cross-Media Joint Representation With Sparse and Semisupervised Regularization

Learning Cross-Media Joint Representation With Sparse and Semisupervised Regularization
复制标题

使用稀疏和半监督正则化学习跨媒体联合表示

DOI:
10.1109/tcsvt.2013.2276704
复制
发表时间:
2014-06-01
影响因子:
8.4
通讯作者:
Xiao, Jianguo
Xiao, Jianguo
中科院分区:
工程技术1区
文献类型:
--
作者:
Zhai, Xiaohua;Peng, Yuxin;Xiao, Jianguo

文献摘要

被引文献

相似文献

跨媒体检索已经成为研究和应用中的一个关键问题,用户可以通过提交任何媒体类型的查询来搜索所有媒体类型(文本、图像、音频、视频和3-D)的结果。如何度量不同媒体之间的内容相似性是关键的挑战。现有的跨媒体检索方法通常侧重于对两两相关或语义信息分别建模。事实上,这两种信息是相辅相成的,同时优化它们可以进一步提高准确性。在本文中,我们提出了一种新的跨媒体数据的特征学习算法,称为联合表示学习(JRL),它能够在一个统一的优化框架中联合探索相关性和语义信息。JRL将不同媒体类型的稀疏和半监督正则化集成到一个统一的优化问题中,而现有的特征学习方法通常专注于单一媒体类型。一方面,JRL同时学习不同介质的稀疏投影矩阵,因此不同介质可以彼此对齐,这对噪声具有鲁棒性。另一方面,对不同媒体类型的标记数据和未标记数据进行了研究。不同媒体类型的未标记示例增加了训练数据的多样性,并提高了联合表示学习的性能。此外,JRL不仅可以降低原始特征的维数,而且可以将跨媒体相关性融入到最终的表示中,从而进一步提高了跨媒体检索和单媒体检索的性能。在两个数据集上的实验表明,与最先进的方法相比,我们所提出的方法的有效性,多达五种媒体类型。
Cross-media retrieval has become a key problem in both research and application, in which users can search results across all of the media types (text, image, audio, video, and 3-D) by submitting a query of any media type. How to measure the content similarity among different media is the key challenge. Existing cross-media retrieval methods usually focus on modeling the pairwise correlation or semantic information separately. In fact, these two kinds of information are complementary to each other and optimizing them simultaneously can further improve the accuracy. In this paper, we propose a novel feature learning algorithm for cross-media data, called joint representation learning (JRL), which is able to explore jointly the correlation and semantic information in a unified optimization framework. JRL integrates the sparse and semisupervised regularization for different media types into one unified optimization problem, while existing feature learning methods generally focus on a single media type. On one hand, JRL learns sparse projection matrix for different media simultaneously, so different media can align with each other, which is robust to the noise. On the other hand, both the labeled data and unlabeled data of different media types are explored. Unlabeled examples of different media types increase the diversity of training data and boost the performance of joint representation learning. Furthermore, JRL can not only reduce the dimension of the original features, but also incorporate the cross-media correlation into the final representation, which further improves the performance of both cross-media retrieval and single-media retrieval. Experiments on two datasets with up to five media types show the effectiveness of our proposed approach, as compared with the state-of-the-art methods.