Block-Row Sparse Multiview Multilabel Learning for Image Classification

Block-Row Sparse Multiview Multilabel Learning for Image Classification
复制标题

用于图像分类的块行稀疏多视图多标签学习

DOI:
10.1109/tcyb.2015.2403356
复制
发表时间:
2016-02-01
影响因子:
11.8
通讯作者:
Zhang, Shichao
Zhang, Shichao
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhu, Xiaofeng;Li, Xuelong;Zhang, Shichao

文献摘要

被引文献

相似文献

在图像分析中,图像通常由多个视觉特征(也称为多视图特征)表示,旨在更好地解释它们,以实现卓越的学习性能。由于每个视图上的特征提取过程是分离的,因此图像的多个视觉特征可能包括重叠、噪声和冗余。因此,使用数据的所有派生视图进行学习可能会降低有效性。为了解决这个问题,本文同时进行了分层特征选择和多视图多标签(MVML)学习多视图图像分类,通过嵌入一个新的块行正则化到MVML框架。将Frobenius范数(F-norm)正则化器和l2,1-norm正则化器连接起来的块行正则化器被设计用于进行分层特征选择,其中F-norm正则化器用于进行高层特征选择以选择信息视图(即,丢弃无信息的视图),然后使用12,1-范数正则化器对信息视图进行低级特征选择。使用块-行正则化器的基本原理是避免过度拟合的问题(通过块-行正则化器),去除冗余视图并保留数据的自然组结构(通过F-范数正则化器),以及分别去除噪声特征(12,1-范数正则化器)。我们进一步设计了一个计算效率高的算法来优化导出的目标函数,并从理论上证明了所提出的优化方法的收敛性。最后,在真实的图像数据集上的实验结果表明,该方法在分类性能上优于两种基线算法和三种最新算法。
In image analysis, the images are often represented by multiple visual features (also known as multiview features), that aim to better interpret them for achieving remarkable performance of the learning. Since the processes of feature extraction on each view are separated, the multiple visual features of images may include overlap, noise, and redundancy. Thus, learning with all the derived views of the data could decrease the effectiveness. To address this, this paper simultaneously conducts a hierarchical feature selection and a multiview multilabel (MVML) learning for multiview image classification, via embedding a proposed a new block-row regularizer into the MVML framework. The block-row regularizer concatenating a Frobenius norm (F-norm) regularizer and an l2,1-norm regularizer is designed to conduct a hierarchical feature selection, in which the F-norm regularizer is used to conduct a high-level feature selection for selecting the informative views (i.e., discarding the uninformative views) and the 12,1-norm regularizer is then used to conduct a low-level feature selection on the informative views. The rationale of the use of a block-row regularizer is to avoid the issue of the over-fitting (via the block-row regularizer), to remove redundant views and to preserve the natural group structures of data (via the F-norm regularizer), and to remove noisy features (the 12,1-norm regularizer), respectively. We further devise a computationally efficient algorithm to optimize the derived objective function and also theoretically prove the convergence of the proposed optimization method. Finally, the results on real image datasets show that the proposed method outperforms two baseline algorithms and three state-of-the-art algorithms in terms of classification performance.