Performance of feature-selection methods in the classification of high-dimension data

Performance of feature-selection methods in the classification of high-dimension data
复制标题

DOI:
10.1016/j.patcog.2008.08.001
复制
发表时间:
2009-03-01
影响因子:
8
通讯作者:
Dougherty, Edward R.
Dougherty, Edward R.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hua, Jianping;Tembe, Waibhav D.;Dougherty, Edward R.

文献摘要

被引文献

相似文献

当代生物技术产生极高维度的数据集来设计分类器,其中有20,000或更多的潜在特征是常见的。此外,样本量往往很小。在这种情况下,特征选择是分类器设计中不可避免的一部分。到目前为止,已经有许多关于特征选择的比较研究,但他们要么考虑的设置比当前生物信息学应用中的设置小得多,要么将他们的研究限制在少数真实数据集上。本研究比较了基于模型的合成数据和真实数据在涉及数千个特征的环境下的几种基本特征选择方法。它定义了涉及不同数量的标记(有用的特征)和非标记(无用的特征)以及特征之间不同类型关系的分布模型。在此框架下,评估了不同分布模型和分类器的特征选择算法的性能。计算分类误差和发现标记的数量。尽管结果清楚地表明,所考虑的特征选择方法中没有一种在所有场景中表现最好,但相对于样本量和特征之间的关系,存在一些总体趋势。例如,与分类器无关的单变量过滤方法也有类似的趋势。对于较难的问题,过滤方法(如t检验)与包装方法具有更好或相似的性能。这种改进的性能通常伴随着显著的峰值。当样本量足够大时,包装方法具有更好的性能。在大多数情况下,与分类器无关的多变量滤波方法ReliefF的性能比单变量滤波方法差;然而,基于relief的包装器方法的性能与基于t测试的包装器方法相似。(C) 2008 Elsevier Ltd版权所有。
Contemporary biological technologies produce extremely high-dimensional data sets from which to design classifiers, with 20,000 or more potential features being common place. In addition, sample sizes tend to be small. In such settings, feature selection is an inevitable part of classifier design. Heretofore, there have been a number of comparative studies for feature selection, but they have either considered settings with much smaller dimensionality than those occurring in current bioinformatics applications or constrained their study to a few real data sets. This study compares some basic feature-selection methods in settings involving thousands of features, using both model-based synthetic data and real data. It defines distribution models involving different numbers of markers (useful features) versus non-markers (useless features) and different kinds of relations among the features. Under this framework, it evaluates the performances of feature-selection algorithms for different distribution models and classifiers. Both classification error and the number of discovered markers are computed. Although the results clearly show that none of the considered feature-selection methods performs best across all scenarios, there are some general trends relative to sample size and relations among the features. For instance, the classifier-independent univariate filter methods have similar trends. Filter methods such as the t-test have better or similar performance with wrapper methods for harder problems. This improved performance is usually accompanied with significant peaking. Wrapper methods have better performance when the sample size is sufficiently large. ReliefF, the classifier-independent multivariate filter method, has worse performance than univariate filter methods in most cases; however, ReliefF-based wrapper methods show performance similar to their t-test-based counterparts. (C) 2008 Elsevier Ltd. All rights reserved.