Impossibility of successful classification when useful features are rare and weak

Impossibility of successful classification when useful features are rare and weak
复制标题

DOI:
10.1073/pnas.0903931106
复制
发表时间:
2009-06-02
影响因子:
11.1
通讯作者:
Jin, Jiashun
Jin, Jiashun
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Jin, Jiashun

文献摘要

被引文献

相似文献

我们研究了一个两类分类问题,有大量的功能,其中许多是无用的,只有少数是有用的,但我们不知道他们是。与训练观察的数量相比,特征的数量很大。用4个关键参数校准模型-特征的数量,训练样本的大小,有用特征的分数和强度-我们在参数空间中确定了一个区域,在这个区域中,没有经过训练的分类器可以可靠地将两个类别在新数据上分开。这个区域的补充,成功的分类是可能的,也简要讨论。
We study a two-class classification problem with a large number of features, out of which many are useless and only a few are useful, but we do not know which ones they are. The number of features is large compared with the number of training observations. Calibrating the model with 4 key parameters-the number of features, the size of the training sample, the fraction, and strength of useful features-we identify a region in parameter space where no trained classifier can reliably separate the two classes on fresh data. The complement of this region-where successful classification is possible-is also briefly discussed.