HIGH-DIMENSIONAL CLASSIFICATION USING FEATURES ANNEALED INDEPENDENCE RULES

HIGH-DIMENSIONAL CLASSIFICATION USING FEATURES ANNEALED INDEPENDENCE RULES
复制标题

DOI:
10.1214/07-aos504
复制
发表时间:
2008-12-01
影响因子:
4.5
通讯作者:
Fan, Yingying
Fan, Yingying
中科院分区:
数学1区
文献类型:
--
作者:
Fan, Jianqing;Fan, Yingying

文献摘要

被引文献

相似文献

使用高维特征的分类经常出现在许多当代统计研究中,例如使用微阵列或其他高通量数据的肿瘤分类。维度对分类的影响知之甚少。在一篇开创性的论文中,Bickel和Levina [Bernoulli 10(2004)989-1010]表明,由于光谱发散,Fisher判别式表现不佳,他们建议使用独立规则来克服这个问题。我们首先证明,即使是独立的分类规则,分类使用所有的功能可以作为穷人的随机猜测,由于在高维特征空间中估计人口质心的噪声积累。事实上,我们进一步证明,几乎所有的线性判别式可以执行的随机猜测一样差。因此,重要的是要选择一个子集的重要特征的高维分类,导致特征退火独立规则(FAIR)。建立了双样本t统计量选择所有重要特征的条件。最佳数量的功能,或等价地,阈值的测试统计量的选择提出的分类误差的上限的基础上。仿真研究和真实的数据分析支持我们的理论结果,并令人信服地证明了我们的新的分类程序的优势。
Classification using high-dimensional features arises frequently in many contemporary statistical studies such as tumor classification using microarray or other high-throughput data. The impact of dimensionality on classifications is poorly understood. In a seminal paper, Bickel and Levina [Bernoulli 10 (2004) 989-1010] show that the Fisher discriminant performs poorly due to diverging spectra and they propose to use the independence rule to overcome the problem. We first demonstrate that even for the independence classification rule, classification using all the features can be as poor as the random guessing due to noise accumulation in estimating population centroids in high-dimensional feature space. In fact, we demonstrate further that almost all linear discriminants can perform as poorly as the random guessing. Thus, it is important to select a subset of important features for high-dimensional classification, resulting in Features Annealed Independence Rules (FAIR). The conditions under which all the important features can be selected by the two-sample t-statistic are established. The choice of the optimal number of features, or equivalently, the threshold value of the test statistics are proposed based on an upper bound of the classification error. Simulation studies and real data analysis support our theoretical results and demonstrate convincingly the advantage of our new classification procedure.