Large-dimensionality small-instance set feature selection: A hybrid bio-inspired heuristic approach

Large-dimensionality small-instance set feature selection: A hybrid bio-inspired heuristic approach
复制标题

DOI:
10.1016/j.swevo.2018.02.021
复制
发表时间:
2018-10
期刊:
Swarm Evol. Comput.
影响因子:
--
通讯作者:
Hossam M. Zawbaa;E. Emary;C. Grosan;V. Snás̃el
Hossam M. Zawbaa;E. Emary;C. Grosan;V. Snás̃el
中科院分区:
其他
文献类型:
--
作者:
Hossam M. Zawbaa;E. Emary;C. Grosan;V. Snás̃el

文献摘要

被引文献

相似文献

在机器学习中,选择具有代表性的特征集仍然是一个至关重要且具有挑战性的问题。当以下任何一种情况发生时,问题的复杂性都会增加:非常大量的属性(大维度);非常少量的实例或时间点(小实例集)。第一种情况给机器学习算法带来了问题,因为用于选择相关特征组合的搜索空间变得不可能在合理的时间内以合理的计算资源进行探索。第二个方面造成的问题是没有足够的数据可供学习(例子不足)。在这项工作中,我们同时处理这两个问题。我们提出的方法是受自然(特别是生物学)启发的方法。我们提出了一种混合的两种方法,其优点是提供了一个很好的学习,从更少的例子和一个公平的选择功能,从一个真正的大集合,所有这些,同时确保高标准的分类精度的数据。使用的方法是蚁群优化(ALO),灰狼优化(GWO),以及两者的组合(ALO-GWO)。我们在具有近50,000个特征和不到200个实例的数据集上测试了它们的性能。与遗传算法(GA)和粒子群优化算法(PSO)等方法相比,该方法具有良好的应用前景。
Selection of a representative set of features is still a crucial and challenging problem in machine learning. The complexity of the problem increases when any of the following situations occur: a very large number of attributes (large dimensionality); a very small number of instances or time points (small-instance set). The first situation poses problems for machine learning algorithm as the search space for selecting a combination of relevant features becomes impossible to explore in a reasonable time and with reasonable computational resources. The second aspect poses the problem of having insufficient data to learn from (insufficient examples). In this work, we approach both these issues at the same time. The methods we proposed are heuristics inspired by nature (in particular, by biology). We propose a hybrid of two methods which has the advantage of providing a good learning from fewer examples and a fair selection of features from a really large set, all these while ensuring a high standard classification accuracy of the data. The methods used are antlion optimization (ALO), grey wolf optimization (GWO), and a combination of the two (ALO-GWO). We test their performance on datasets having almost 50,000 features and less than 200 instances. The results look promising while compared with other methods such as genetic algorithms (GA) and particle swarm optimization (PSO).