Benchmark for filter methods for feature selection in high-dimensional classification data

Benchmark for filter methods for feature selection in high-dimensional classification data
复制标题

DOI:
10.1016/j.csda.2019.106839
复制
发表时间:
2020-03-01
影响因子:
1.8
通讯作者:
Lang, Michel
Lang, Michel
中科院分区:
数学3区
文献类型:
--
作者:
Bommert, Andrea;Sun, Xudong;Lang, Michel

文献摘要

被引文献

相似文献

特征选择是机器学习中最基本的问题之一,由于生物信息学等不同领域出现的高维数据集,特征选择越来越受到人们的关注。对于特征选择,过滤方法起着重要的作用,因为它们可以与任何机器学习模型相结合,并且可以大大减少机器学习算法的运行时间。分析的目的是审查不同的过滤方法的工作原理,比较它们的性能方面的运行时间和预测精度,并为应用程序提供指导。基于16个高维分类数据集,分析了22种过滤方法与分类方法相结合时的运行时间和精度。它的结论是,有没有一组的过滤器的方法,总是优于所有其他方法,但建议的过滤器的方法,许多数据集上表现良好。此外,还发现了在对特征进行排序的顺序方面相似的过滤器组。为了进行分析,使用了R机器学习包mlr。它提供了一个统一的编程API,因此是一个方便的工具,进行特征选择使用过滤器的方法。(C)2019年,任作家。由爱思唯尔公司出版
Feature selection is one of the most fundamental problems in machine learning and has drawn increasing attention due to high-dimensional data sets emerging from different fields like bioinformatics. For feature selection, filter methods play an important role, since they can be combined with any machine learning model and can heavily reduce run time of machine learning algorithms. The aim of the analyses is to review how different filter methods work, to compare their performance with respect to both run time and predictive accuracy, and to provide guidance for applications. Based on 16 high-dimensional classification data sets, 22 filter methods are analyzed with respect to run time and accuracy when combined with a classification method. It is concluded that there is no group of filter methods that always outperforms all other methods, but recommendations on filter methods that perform well on many of the data sets are made. Also, groups of filters that are similar with respect to the order in which they rank the features are found. For the analyses, the R machine learning package mlr is used. It provides a uniform programming API and therefore is a convenient tool to conduct feature selection using filter methods. (C) 2019 The Authors. Published by Elsevier B.V.