LLE Score: A New Filter-Based Unsupervised Feature Selection Method Based on Nonlinear Manifold Embedding and Its Application to Image Recognition

LLE Score: A New Filter-Based Unsupervised Feature Selection Method Based on Nonlinear Manifold Embedding and Its Application to Image Recognition
复制标题

LLE Score:一种基于非线性流形嵌入的新型滤波器无监督特征选择方法及其在图像识别中的应用

DOI:
10.1109/tip.2017.2733200
复制
发表时间:
2017
影响因子:
10.6
通讯作者:
Han Junwei
Han Junwei
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yao Chao;Liu Ya-Feng;Jiang Bo;Han Jungong;Han Junwei

文献摘要

被引文献

相似文献

特征选择的任务是从原始高维数据中找出最具代表性的特征。由于缺乏类标签的信息,在无监督学习场景中选择合适的特征比在有监督学习场景中要困难得多。本文研究了一种流行的流形学习方法局部线性嵌入(LLE)在特征选择任务中的潜力。将LLE的思想应用于保图特征选择框架是很简单的。然而,我们发现这个简单的应用程序存在一些问题。例如,当特征中的元素都相等时,它就会失败;它不具有尺度不变性,不能有效地捕捉图的变化。为了解决这些问题,本文提出了一种新的基于LLE的基于过滤器的特征选择方法,并将其命名为LLE评分。该准则衡量每个特征的局部结构与原始数据的局部结构之间的差异。我们在两个人脸图像数据集、一个物体图像数据集和一个手写数字数据集上的分类任务实验表明,LLE分数优于最先进的方法,包括数据方差、拉普拉斯分数和稀疏性分数。
The task of feature selection is to find the most representative features from the original high-dimensional data. Because of the absence of the information of class labels, selecting the appropriate features in unsupervised learning scenarios is much harder than that in supervised scenarios. In this paper, we investigate the potential of locally linear embedding (LLE), which is a popular manifold learning method, in feature selection task. It is straightforward to apply the idea of LLE to the graph-preserving feature selection framework. However, we find that this straightforward application suffers from some problems. For example, it fails when the elements in the feature are all equal; it does not enjoy the property of scaling invariance and cannot capture the change of the graph efficiently. To solve these problems, we propose a new filter-based feature selection method based on LLE in this paper, which is named as LLE score. The proposed criterion measures the difference between the local structure of each feature and that of the original data. Our experiments of classification task on two face image data sets, an object image data set, and a handwriting digits data set show that LLE score outperforms state-of-the-art methods, including data variance, Laplacian score, and sparsity score.