Ensemble feature selection for high-dimensional data: a stability analysis across multiple domains

Ensemble feature selection for high-dimensional data: a stability analysis across multiple domains
复制标题

DOI:
10.1007/s00521-019-04082-3
复制
发表时间:
2020-05-01
影响因子:
6
通讯作者:
Pes, Barbara
Pes, Barbara
中科院分区:
计算机科学3区
文献类型:
--
作者:
Pes, Barbara

文献摘要

被引文献

相似文献

选择相关特征的子集对于分析来自许多应用领域(如生物医学数据、文档和图像分析)的高维数据集至关重要。由于没有单一的选择算法能够在预测性能和稳定性(即对输入数据变化的鲁棒性)方面确保最佳结果,研究人员越来越多地探索涉及不同选择器组合的“集成”方法的有效性。虽然文献中已经报道了一些有趣的建议,但到目前为止,它们中的大多数都是在有限的设置中进行评估的(例如,使用来自单个领域的数据并与特定的选择方法相结合),关于集成特征选择的大规模适用性和实用性的重要问题没有得到回答。为了对该领域做出贡献,这项工作提出了一项实证研究,该研究涵盖了不同类型的选择算法(过滤器和嵌入式方法,单变量和多变量技术)和不同的应用领域。具体来说,我们考虑了18个具有异构特征的分类任务(就类的数量和实例与特征的比率而言),并通过实验评估了不同基数的特征子集,在多大程度上集成方法比单个选择器更健壮,从而为研究人员和从业者提供了有用的见解。
Selecting a subset of relevant features is crucial to the analysis of high-dimensional datasets coming from a number of application domains, such as biomedical data, document and image analysis. Since no single selection algorithm seems to be capable of ensuring optimal results in terms of both predictive performance and stability (i.e. robustness to changes in the input data), researchers have increasingly explored the effectiveness of "ensemble" approaches involving the combination of different selectors. While interesting proposals have been reported in the literature, most of them have been so far evaluated in a limited number of settings (e.g. with data from a single domain and in conjunction with specific selection approaches), leaving unanswered important questions about the large-scale applicability and utility of ensemble feature selection. To give a contribution to the field, this work presents an empirical study which encompasses different kinds of selection algorithms (filters and embedded methods, univariate and multivariate techniques) and different application domains. Specifically, we consider 18 classification tasks with heterogeneous characteristics (in terms of number of classes and instances-to-features ratio) and experimentally evaluate, for feature subsets of different cardinalities, the extent to which an ensemble approach turns out to be more robust than a single selector, thus providing useful insight for both researchers and practitioners.