INSTANCE-BASED LEARNING ALGORITHMS

INSTANCE-BASED LEARNING ALGORITHMS
复制标题

DOI:
10.1007/bf00153759
复制
发表时间:
1991-01-01
期刊:
影响因子:
7.5
通讯作者:
ALBERT, MK
ALBERT, MK
中科院分区:
计算机科学3区
文献类型:
--
作者:
AHA, DW;KIBLER, D;ALBERT, MK

文献摘要

被引文献

相似文献

存储和使用特定实例可以提高几种监督学习算法的性能。 这些包括学习决策树、分类规则和分布式网络的算法。 然而,没有调查分析的算法,只使用特定的实例来解决增量学习任务。 在本文中,我们描述了一个框架和方法,称为基于实例的学习,只使用特定的实例生成分类预测。 基于实例的学习算法不维护从特定实例导出的一组抽象。 该方法扩展了最近邻算法,该算法具有较大的存储需求。 我们描述了如何存储需求可以显着减少,最多,在学习率和分类精度的小牺牲。 虽然存储减少算法在几个真实世界的数据库上表现良好,但其性能随着训练实例中的属性噪声水平而迅速下降。 因此,我们用显著性检验来区分噪声实例。 这种扩展算法的性能随着噪声水平的增加而优雅地下降,并且与噪声容忍决策树算法相比毫不逊色。
Storing and using specific instances improves the performance of several supervised learning algorithms. These include algorithms that learn decision trees, classification rules, and distributed networks. However, no investigation has analyzed algorithms that use only specific instances to solve incremental learning tasks. In this paper, we describe a framework and methodology, called instance-based learning, that generates classification predictions using only specific instances. Instance-based learning algorithms do not maintain a set of abstractions derived from specific instances. This approach extends the nearest neighbor algorithm, which has large storage requirements. We describe how storage requirements can be significantly reduced with, at most, minor sacrifices in learning rate and classification accuracy. While the storage-reducing algorithm performs well on several real-world databases, its performance degrades rapidly with the level of attribute noise in training instances. Therefore, we extended it with a significance test to distinguish noisy instances. This extended algorithm's performance degrades gracefully with increasing noise levels and compares favorably with a noise-tolerant decision tree algorithm.