NEAREST NEIGHBOR PATTERN CLASSIFICATION

NEAREST NEIGHBOR PATTERN CLASSIFICATION
复制标题

DOI:
10.1109/tit.1967.1053964
复制
发表时间:
1967-01-01
影响因子:
2.5
通讯作者:
HART, PE
HART, PE
中科院分区:
计算机科学2区
文献类型:
--
作者:
COVER, TM;HART, PE

文献摘要

被引文献

相似文献

最近邻判决规则将一组先前已分类的点中的最近点的分类分配给未分类的样本点。这一规则与样本点及其分类的基本联合分布无关,因此,这种规则的错误概率必须至少与贝叶斯错误概率一样大--考虑到基本概率结构的所有决策规则的最小错误概率。然而,在大样本分析中,我们将在-CATEGORY的情况下表明,对于所有适当平滑的基础分布,这些界限是最紧的。因此,对于任意数量的类别,最近邻规则的错误概率被限定为贝叶斯错误概率的两倍以上。在这个意义上,可以说无限样本集中一半的分类信息包含在最近的邻居中。
The nearest neighbor decision rule assigns to an unclassified sample point the classification of the nearest of a set of previously classified points. This rule is independent of the underlying joint distribution on the sample points and their classifications, and hence the probability of errorof such a rule must be at least as great as the Bayes probability of error--the minimum probability of error over all decision rules taking underlying probability structure into account. However, in a large sample analysis, we will show in the-category case that, where these bounds are the tightest possible, for all suitably smooth underlying distributions. Thus for any number of categories, the probability of error of the nearest neighbor rule is bounded above by twice the Bayes probability of error. In this sense, it may be said that half the classification information in an infinite sample set is contained in the nearest neighbor.