Integer programming models for feature selection: New extensions and a randomized solution algorithm

Integer programming models for feature selection: New extensions and a randomized solution algorithm
复制标题

DOI:
10.1016/j.ejor.2015.09.051
复制
发表时间:
2016-04-16
影响因子:
6.4
通讯作者:
Weitschek, E.
Weitschek, E.
中科院分区:
管理学2区
文献类型:
--
作者:
Bertolazzi, P.;Felici, G.;Weitschek, E.

文献摘要

被引文献

相似文献

特征选择方法用于机器学习和数据分析中,以选择可以成功用于构建数据模型的特征子集。这些方法是在假设许多可用特征对于分析目的而言是冗余的情况下应用的。在本文中,我们专注于一个特殊的方法,在监督学习问题的特征选择,基于一个线性,规划模型与整数变量。对于与这种方法相关的优化问题的解决方案,我们提出了一种新的强大的元算法,它依赖于一个贪婪的随机自适应搜索过程,通过短内存和局部搜索策略扩展。我们的启发式算法的性能成功地比较了那些成熟的特征选择方法,无论是模拟和真实的数据从生物应用。所获得的结果表明,我们的方法是特别适合于具有非常大量的二进制或分类功能的问题。(C)2015年,Elsevier B.V.和欧洲运筹学会协会(EURO)在国际运筹学会联合会(IFORS)内。All rights reserved.
Feature selection methods are used in machine learning and data analysis to select a subset of features that may be successfully used in the construction of a model for the data. These methods are applied under the assumption that often many of the available features are redundant for the purpose of the analysis. In this paper, we focus on a particular method for feature selection in supervised learning problems, based on a linear, programming model with integer variables. For the solution of the optimization problem associated with this approach, we propose a novel robust metaheuristics algorithm that relies on a Greedy Randomized Adaptive Search Procedure, extended with the adoption of short memory and a local search strategy. The performances of our heuristic algorithm are successfully compared with those of well-established feature selection methods, both on simulated and real data from biological applications. The obtained results suggest that our method is particularly suited for problems with a very large number of binary or categorical features. (C) 2015 Elsevier B.V. and Association of European Operational Research Societies (EURO) within the International Federation of Operational Research Societies (IFORS). All rights reserved.