Trace Ratio Criterion for Feature Selection

Trace Ratio Criterion for Feature Selection
复制标题

DOI:
--
复制
发表时间:
2008-07
期刊:
--
影响因子:
--
通讯作者:
F. Nie;Shiming Xiang;Yangqing Jia;Changshui Zhang;Shuicheng Yan
F. Nie;Shiming Xiang;Yangqing Jia;Changshui Zhang;Shuicheng Yan
中科院分区:
其他
文献类型:
--
作者:
F. Nie;Shiming Xiang;Yangqing Jia;Changshui Zhang;Shuicheng Yan

文献摘要

被引文献

相似文献

Fisher Score和Laplace Score是两种流行的特征选择算法,它们都属于通用的基于图的特征选择框架。在该框架中,基于相应的分数(子集级别分数)选择特征子集,该分数以跟踪比率的形式计算。由于所有可能的特征子集的数量非常巨大,以暴力方式搜索具有最大子集级别分数的特征子集通常在计算成本上令人望而却步。传统的方法不是计算所有特征子集的得分,而是计算每个特征的得分,然后根据这些特征级别得分的排名来选择领先的特征。然而,基于特征级别得分选择特征子集并不能保证子集级别得分的最优。在本文中,我们直接对子集级别的得分进行优化,并提出了一种新的算法来高效地寻找全局最优的特征子集,使子集级别的得分最大化。大量的实验表明,与传统的特征选择方法相比,本文提出的算法是有效的。
Fisher score and Laplacian score are two popular feature selection algorithms, both of which belong to the general graph-based feature selection framework. In this framework, a feature subset is selected based on the corresponding score (subset-level score), which is calculated in a trace ratio form. Since the number of all possible feature subsets is very huge, it is often prohibitively expensive in computational cost to search in a brute force manner for the feature subset with the maximum subset-level score. Instead of calculating the scores of all the feature subsets, traditional methods calculate the score for each feature, and then select the leading features based on the rank of these feature-level scores. However, selecting the feature subset based on the feature-level score cannot guarantee the optimum of the subset-level score. In this paper, we directly optimize the subset-level score, and propose a novel algorithm to efficiently find the global optimal feature subset such that the subset-level score is maximized. Extensive experiments demonstrate the effectiveness of our proposed algorithm in comparison with the traditional methods for feature selection.