Improving the Performance of Feature Selection Methods with Low-Sample-Size Data

Improving the Performance of Feature Selection Methods with Low-Sample-Size Data
复制标题

DOI:
10.1093/comjnl/bxac033
复制
发表时间:
2022-04
期刊:
Comput. J.
影响因子:
--
通讯作者:
Wanwan Zheng;Mingzhe Jin
Wanwan Zheng;Mingzhe Jin
中科院分区:
其他
文献类型:
--
作者:
Wanwan Zheng;Mingzhe Jin

文献摘要

相似文献

特征选择是指机器学习的关键预处理,以去除不相关和冗余的数据。根据特征选择方法,通常需要足够的样本来选择可靠的特征子集,特别是考虑到离群值的存在。然而,在一些现实世界的应用中(例如神经科学,生物信息学和心理学),不能总是确保足够的样本。提出了一种基于数据质量和可变训练样本的特征选择方法(QVT),以提高超低样本数据下特征选择方法的性能。鉴于没有一种特征选择方法可以在所有场景中实现最佳性能,QVT的主要特点是其多功能性,因为它可以在任何特征选择方法中实现。此外,与现有的通过增加样本量或使用更复杂的算法来提取低样本量数据的稳定特征子集的方法相比,QVT试图利用原始数据进行改进。采用20个基准数据集、3种特征选择方法和3种分类器进行实验,验证了QVT的可行性;实验结果表明,使用QVT选择的特征能够获得比使用显式特征选择方法更高的分类精度,且存在显著差异。
Feature selection refers to a critical preprocessing of machine learning to remove irrelevant and redundant data. According to feature selection methods, sufficient samples are usually required to select a reliable feature subset, especially considering the presence of outliers. However, sufficient samples cannot always be ensured in several real-world applications (e.g. neuroscience, bioinformatics and psychology). This study proposed a method to improve the performance of feature selection methods with ultra low-sample-size data, which is named feature selection based on data quality and variable training samples (QVT). Given that none of feature selection methods can perform optimally in all scenarios, QVT is primarily characterized by its versatility, because it can be implemented in any feature selection method. Furthermore, compared to the existing methods which tried to extract a stable feature subset for low-sample-size data by increasing the sample size or using more complicated algorithm, QVT tried to get improvement using the original data. An experiment was performed using 20 benchmark datasets, three feature selection methods and three classifiers to verify the feasibility of QVT; the results showed that using features selected by QVT is capable of achieving higher classification accuracy than using the explicit feature selection method, and significant differences exist.