Feature selection based on artificial bee colony and gradient boosting decision tree

Feature selection based on artificial bee colony and gradient boosting decision tree
复制标题

基于人工蜂群和梯度提升决策树的特征选择

DOI:
10.1016/j.asoc.2018.10.036
复制
发表时间:
2019-01-01
影响因子:
8.7
通讯作者:
Gu, Lichuan
Gu, Lichuan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Rao, Haidi;Shi, Xianzhang;Gu, Lichuan

文献摘要

被引文献

相似文献

来自许多实际应用的数据可能是高维的,并且此类数据的特征通常是高度冗余的。识别信息特征已成为数据挖掘的一个重要步骤,这不仅是为了规避维度灾难,也是为了减少待处理的数据量。在本文中,我们提出了一种基于蜂群和梯度提升决策树的新型特征选择方法,旨在解决所选特征的效率和信息质量等问题。我们的方法利用蜂群算法实现决策树输入的全局优化,以识别信息特征。该方法初始化由数据集所跨越的特征空间。根据使用人工蜂群算法在决策过程中所贡献的信息,不太相关的特征被抑制。我们使用两个乳腺癌数据集以及来自公共数据存储库的六个数据集进行了实验。实验结果表明,所提出的方法有效地降低了数据集的维度,并利用所选特征实现了更高的分类准确率。© 2018爱思唯尔有限公司。保留所有权利。
Data from many real-world applications can be high dimensional and features of such data are usually highly redundant. Identifying informative features has become an important step for data mining to not only circumvent the curse of dimensionality but to reduce the amount of data for processing. In this paper, we propose a novel feature selection method based on bee colony and gradient boosting decision tree aiming at addressing problems such as efficiency and informative quality of the selected features. Our method achieves global optimization of the inputs of the decision tree using the bee colony algorithm to identify the informative features. The method initializes the feature space spanned by the dataset. Less relevant features are suppressed according to the information they contribute to the decision making using an artificial bee colony algorithm. Experiments are conducted with two breast cancer datasets and six datasets from the public data repository. Experimental results demonstrate that the proposed method effectively reduces the dimensions of the dataset and achieves superior classification accuracy using the selected features. (C) 2018 Elsevier B.V. All rights reserved.