Feature selection via binary simultaneous perturbation stochastic approximation.

Feature selection via binary simultaneous perturbation stochastic approximation.
复制标题

DOI:
10.1016/j.patrec.2016.03.002
复制
发表时间:
2016-05-01
影响因子:
5.1
通讯作者:
Malekipirbazari, Milad
Malekipirbazari, Milad
中科院分区:
计算机科学3区
文献类型:
--
作者:
Aksakalli, Vural;Malekipirbazari, Milad

文献摘要

被引文献

相似文献

特征选择(FS)已成为处理当今高度复杂的模式识别问题的大量特征的一个不可或缺的任务。在这项研究中,我们提出了一个新的包装方法FS的基础上二进制同时扰动随机近似(BSPSA)。这种伪梯度下降随机算法从初始特征向量开始,并通过连续迭代向最佳特征向量移动。在每次迭代中,当前特征向量的各个分量同时受到来自合格概率分布的随机偏移的扰动。我们目前的计算实验数据集的功能数量从几十到数千使用三个广泛使用的分类器作为包装:最近邻,决策树和线性支持向量机。我们比较我们的方法对全套功能,以及二进制遗传算法和顺序FS方法使用交叉验证的分类错误率和AUC作为性能标准。我们的研究结果表明,BSPSA选择的功能相比,一般的替代方法和BSPSA可以产生上级功能集的数据集与成千上万的功能,通过检查一个非常小的部分的解决方案空间。我们不知道任何其他包装FS方法,计算上可行的,具有良好的收敛性能,这样的大数据集。(C)© 2016 Elsevier B.V.版权所有。
Feature selection (FS) has become an indispensable task in dealing with today's highly complex pattern recognition problems with massive number of features. In this study, we propose a new wrapper approach for FS based on binary simultaneous perturbation stochastic approximation (BSPSA). This pseudo-gradient descent stochastic algorithm starts with an initial feature vector and moves toward the optimal feature vector via successive iterations. In each iteration, the current feature vector's individual components are perturbed simultaneously by random offsets from a qualified probability distribution. We present computational experiments on datasets with numbers of features ranging from a few dozens to thousands using three widely-used classifiers as wrappers: nearest neighbor, decision tree, and linear support vector machine. We compare our methodology against the full set of features as well as a binary genetic algorithm and sequential FS methods using cross-validated classification error rate and AUC as the performance criteria. Our results indicate that features selected by BSPSA compare favorably to alternative methods in general and BSPSA can yield superior feature sets for datasets with tens of thousands of features by examining an extremely small fraction of the solution space. We are not aware of any other wrapper FS methods that are computationally feasible with good convergence properties for such large datasets. (C) 2016 Elsevier B.V. All rights reserved.