Selecting critical features for data classification based on machine learning methods

Selecting critical features for data classification based on machine learning methods
复制标题

DOI:
10.1186/s40537-020-00327-4
复制
发表时间:
2020-07-23
影响因子:
8.1
通讯作者:
Caraka, Rezzy Eko
Caraka, Rezzy Eko
中科院分区:
计算机科学2区
文献类型:
--
作者:
Chen, Rung-Ching;Dewi, Christine;Caraka, Rezzy Eko

文献摘要

被引文献

相似文献

特征选择变得突出,特别是在具有许多变量和特征的数据集中。它将消除不重要的变量,提高精度以及分类的性能。随机森林已经成为一个非常有用的算法,即使有更多的变量也可以处理特征选择问题。在本文中,我们使用了三个流行的数据集与更多的变量(银行营销,汽车评估数据库,人类活动识别使用智能手机)进行实验。有四个主要原因说明为什么特征选择是必不可少的。首先,通过减少参数的数量来简化模型,然后减少训练时间,通过增强泛化能力来减少过度填充,并避免维数灾难。此外,我们还评估和比较了随机森林(RF),支持向量机(SVM),K-最近邻(KNN)和线性判别分析(LDA)等分类模型的准确性和性能。准确率最高的模型就是最好的分类器。在实际应用中,本文采用随机森林算法来选择重要特征。我们的实验清楚地显示了RF算法从不同角度的比较研究。此外,我们还比较了使用RF方法varImp()、Boruta和递归特征消除(RFE)进行基本特征选择和不进行基本特征选择的数据集的结果,以获得最佳的准确率和kappa。实验结果表明,随机森林在所有实验组中均取得了较好的性能。
Feature selection becomes prominent, especially in the data sets with many variables and features. It will eliminate unimportant variables and improve the accuracy as well as the performance of classification. Random Forest has emerged as a quite useful algorithm that can handle the feature selection issue even with a higher number of variables. In this paper, we use three popular datasets with a higher number of variables (Bank Marketing, Car Evaluation Database, Human Activity Recognition Using Smartphones) to conduct the experiment. There are four main reasons why feature selection is essential. First, to simplify the model by reducing the number of parameters, next to decrease the training time, to reduce overfilling by enhancing generalization, and to avoid the curse of dimensionality. Besides, we evaluate and compare each accuracy and performance of the classification model, such as Random Forest (RF), Support Vector Machines (SVM), K-Nearest Neighbors (KNN), and Linear Discriminant Analysis (LDA). The highest accuracy of the model is the best classifier. Practically, this paper adopts Random Forest to select the important feature in classification. Our experiments clearly show the comparative study of the RF algorithm from different perspectives. Furthermore, we compare the result of the dataset with and without essential features selection by RF methods varImp(), Boruta, and Recursive Feature Elimination (RFE) to get the best percentage accuracy and kappa. Experimental results demonstrate that Random Forest achieves a better performance in all experiment groups.