Constraint Score: A new filter method for feature selection with pairwise constraints

Constraint Score: A new filter method for feature selection with pairwise constraints
复制标题

DOI:
10.1016/j.patcog.2007.10.009
复制
发表时间:
2008-05
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
Daoqiang Zhang;Songcan Chen;Zhi-Hua Zhou
Daoqiang Zhang;Songcan Chen;Zhi-Hua Zhou
中科院分区:
其他
文献类型:
--
作者:
Daoqiang Zhang;Songcan Chen;Zhi-Hua Zhou

文献摘要

被引文献

相似文献

特征选择是挖掘高维数据的重要预处理步骤。一般来说,有监督信息的有监督特征选择方法优于没有监督信息的无监督特征选择方法。在文献中,几乎所有现有的有监督特征选择方法都使用类标签作为监督信息。在本文中,我们建议使用另一种形式的监督信息进行特征选择,即成对约束,它指定一对数据样本属于同一类(必须链接约束)还是不同类(不能链接约束)。成对约束在许多任务中自然出现,并且比类标签更实用、更便宜。该主题尚未在特征选择研究中得到解决。我们将成对约束引导特征选择算法称为约束分数,并将其与著名的费舍尔分数和拉普拉斯分数算法进行比较。在几个高维UCI和人脸数据集上进行了实验。实验结果表明,在很少的成对约束条件下,Constraint Score 在整个训练数据上取得了与带有全类标签的 Fisher Score 相似甚至更高的性能,并且显着优于 Laplacian Score。
Feature selection is an important preprocessing step in mining high-dimensional data. Generally, supervised feature selection methods with supervision information are superior to unsupervised ones without supervision information. In the literature, nearly all existing supervised feature selection methods use class labels as supervision information. In this paper, we propose to use another form of supervision information for feature selection, i.e. pairwise constraints, which specifies whether a pair of data samples belong to the same class (must-link constraints) or different classes (cannot-link constraints). Pairwise constraints arise naturally in many tasks and are more practical and inexpensive than class labels. This topic has not yet been addressed in feature selection research. We call our pairwise constraints guided feature selection algorithm as Constraint Score and compare it with the well-known Fisher Score and Laplacian Score algorithms. Experiments are carried out on several high-dimensional UCI and face data sets. Experimental results show that, with very few pairwise constraints, Constraint Score achieves similar or even higher performance than Fisher Score with full class labels on the whole training data, and significantly outperforms Laplacian Score.