High-dimensional Classification Using Features Annealed Independence Rules 1

High-dimensional Classification Using Features Annealed Independence Rules 1
复制标题

使用特征退火独立规则 1 的高维分类

DOI:
--
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
Yingying Fan
Yingying Fan
中科院分区:
--
文献类型:
--
作者:
Jianqing Fan;Yingying Fan

文献摘要

被引文献

相似文献

使用高维特征进行分类经常出现在许多当代统计学研究中,例如使用微阵列或其他高通量数据进行肿瘤分类。维度对分类的影响还知之甚少。在一篇开创性的论文中,Bickel和Levina[Bernoulli 10(2004)989-1010]证明了由于光谱发散,Fisher判别式的表现很差,他们建议使用独立性规则来克服这个问题。我们首先证明了即使对于独立分类规则,在高维特征空间中估计种群质心时,使用所有特征的分类也可能与由于噪声积累而导致的随机猜测一样差。事实上,我们进一步证明了几乎所有的线性判别式都可以和随机猜测一样差。因此,重要的是选择重要特征的子集进行高维分类,从而产生特征退火式独立规则(FAIR)。建立了用两样本t统计量选择所有重要特征的条件。基于分类误差的上界,提出了最优特征数的选择,即测试统计量的阈值。仿真研究和实际数据分析支持了我们的理论结果,并令人信服地证明了我们新的分类方法的优越性。1.引言。随着成像技术的快速发展,高通量数据如微阵列和蛋白质组学数据经常出现在许多当代统计研究中。例如,在微阵列数据的分析中,维度往往是数千或更多,而样本
Classification using high-dimensional features arises frequently in many contemporary statistical studies such as tumor classification using microarray or other high-throughput data. The impact of dimensionality on classifications is poorly understood. In a seminal paper, Bickel and Levina [Bernoulli 10 (2004) 989–1010] show that the Fisher discriminant performs poorly due to diverging spectra and they propose to use the independence rule to overcome the problem. We first demonstrate that even for the independence classification rule, classification using all the features can be as poor as the random guessing due to noise accumulation in estimating population centroids in high-dimensional feature space. In fact, we demonstrate further that almost all linear discriminants can perform as poorly as the random guessing. Thus, it is important to select a subset of important features for high-dimensional classification, resulting in Features Annealed Independence Rules (FAIR). The conditions under which all the important features can be selected by the two-sample t-statistic are established. The choice of the optimal number of features , or equivalently, the threshold value of the test statistics are proposed based on an upper bound of the classification error. Simulation studies and real data analysis support our theoretical results and demonstrate convincingly the advantage of our new classification procedure. 1. Introduction. With rapid advance of imaging technology, high-through-put data such as microarray and proteomics data are frequently seen in many contemporary statistical studies. For instance, in the analysis of Microarray data, the dimensionality is frequently thousands or more, while the sample