Multi-label feature selection with missing labels

Multi-label feature selection with missing labels
复制标题

DOI:
10.1016/j.patcog.2017.09.036
复制
发表时间:
2018-02
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
Peng Fei Zhu;Qian Xu;Q. Hu;Changqing Zhang;Hong Zhao
Peng Fei Zhu;Qian Xu;Q. Hu;Changqing Zhang;Hong Zhao
中科院分区:
其他
文献类型:
--
作者:
Peng Fei Zhu;Qian Xu;Q. Hu;Changqing Zhang;Hong Zhao

文献摘要

被引文献

相似文献

特征维数的不断增加给多标签学习带来了巨大的时间复杂度和存储负担。为了减轻高维特征的影响,开发了许多多标签特征选择技术。现有的多标签特征选择算法假设训练数据的标签是完整的。然而,这个假设并不总是正确的,因为标记数据是昂贵的,而且类之间存在歧义。因此,在实际应用程序中,可用的数据通常具有一组不完整的标签。本文提出了一种新的标签缺失情况下的多标签特征选择模型。该算法可以选择最具判别性的特征,同时恢复缺失的标签。为了去除不相关和有噪声的特征,对特征选择矩阵施加有效的2,p-范数(0 <p≤ 1)正则化项。为了解决优化问题,我们提出了一种保证收敛的迭代加权最小二乘算法。在基准数据集上的实验结果表明,该方法优于目前最先进的多标签特征选择算法。
The consistently increasing of the feature dimension brings about great time complexity and storage burden for multi-label learning. Numerous multi-label feature selection techniques are developed to alleviate the effect of high-dimensionality. The existing multi-label feature selection algorithms assume that the labels of the training data are complete. However, this assumption does not always hold true for labeling data is costly and there is ambiguity among classes. Hence, in real-world applications, the data available usually have an incomplete set of labels. In this paper, we present a novel multi-label feature selection model under the circumstance of missing labels. With the proposed algorithm, the most discriminative features are selected and missing labels are recovered simultaneously. To remove the irrelevant and noisy features, the effectivel2,p-norm (0 <p≤ 1) regularization item is imposed on the feature selection matrix. To solve the optimization problem, we developed an iterative reweighted least squares (IRLS) algorithm with guaranteed convergence. Experimental results on benchmark datasets show that the proposed method outperforms the state-of-the-art multi-label feature selection algorithms.