Trace Ratio Criterion Based Large Margin Subspace Learning for Feature Selection

Trace Ratio Criterion Based Large Margin Subspace Learning for Feature Selection
复制标题

基于迹比准则的大裕度子空间学习用于特征选择

DOI:
10.1109/access.2018.2888924
复制
发表时间:
2019
期刊:
影响因子:
3.9
通讯作者:
Han Jiqing
Han Jiqing
中科院分区:
计算机科学3区
文献类型:
--
作者:
Luo Hui;Han Jiqing

文献摘要

参考文献

相似文献

在本文中,我们提出了一种基于子空间学习并使用大余量原理的新颖特征选择模型。首先,我们提出一个由给定实例及其最近缺失和最近命中描述的新边缘度量,这可以分别解释为具有不同标签和相同标签的最近邻居。具体来说,对于给定的实例,边距是最近缺失的距离与最近命中的距离的比率,而不是距离差,这有助于更好的平衡,因为到最近缺失的距离通常比最近命中的距离大得多。所提出的模型寻找一个子空间,其中边际度量最大化。此外,考虑到给定样本的最近邻在存在许多不相关特征的情况下是不确定的,我们将它们视为隐藏变量并估计边际的期望。为了执行特征选择,对子空间投影矩阵施加 $\ell _{2,1}$ -范数以强制行稀疏性。由此产生的迹线比率优化问题可以与非线性特征值问题联系起来,很难解决。因此,我们设计了一种有效的迭代算法,并对收敛性进行了理论分析。最后,我们通过与其他几种最先进的方法进行比较来评估所提出的方法。对现实世界数据集的广泛实验表明了我们提出的方法的优越性。
In this paper, we propose a novel feature selection model based on subspace learning with the use of a large margin principle. First, we present a new margin metric described by a given instance and its nearest missing and nearest hit, which can be explained as the nearest neighbor with a different label and the same label, respectively. Specifically, for a given instance, the margin is the ratio of the distance of the nearest missing to that of the nearest hit rather than the difference of distances, which contributes to better balance since the distance to the nearest missing is usually much larger than the nearest hit. The proposed model seeks a subspace in which the margin metric is maximized. Moreover, considering that the nearest neighbors of a given sample are uncertain in the presence of many irrelevant features, we treat them as hidden variables and estimate the expectation of margin. To perform the feature selection, an $\ell _{2,1}$ -norm is imposed on the subspace projection matrix to enforce row sparsity. The resulting trace ratio optimization problem, which can be connected to a nonlinear eigenvalue problem, is hard to solve. Thus, we design an efficient iterative algorithm and present a theoretical analysis of the convergence. Finally, we evaluate the proposed method by comparing it against several other state-of-the-art methods. The extensive experiments on real-world datasets show the superiority of our proposed approach.
DOI: --
发表时间: 2008-07
期刊: --
影响因子: --
作者:
F. Nie;Shiming Xiang;Yangqing Jia;Changshui Zhang;Shuicheng Yan
通讯作者: F. Nie;Shiming Xiang;Yangqing Jia;Changshui Zhang;Shuicheng Yan
DOI: 10.1109/cvpr.2007.382983
发表时间: 2007-06
期刊: 2007 IEEE Conference on Computer Vision and Pattern Recognition
影响因子: --
作者:
Huan Wang;Shuicheng Yan;Dong Xu;Xiaoou Tang;Thomas S. Huang
通讯作者: Huan Wang;Shuicheng Yan;Dong Xu;Xiaoou Tang;Thomas S. Huang
DOI: --
发表时间: 2002
期刊: --
影响因子: --
作者:
K. Crammer;Ran Gilad-Bachrach;A. Navot;Naftali Tishby
通讯作者: K. Crammer;Ran Gilad-Bachrach;A. Navot;Naftali Tishby
DOI: 10.1109/tpami.2007.250598
发表时间: 2007-01-01
影响因子: 23.6
作者:
Yan, Shuicheng;Xu, Dong;Lin, Stephen
通讯作者: Lin, Stephen
DOI: 10.1137/090776603
发表时间: 2010-01-01
影响因子: 1.5
作者:
Ngo, T. T.;Bellalij, M.;Saad, Y.
通讯作者: Saad, Y.