Cluster-based instance selection for machine classification

Cluster-based instance selection for machine classification
复制标题

DOI:
10.1007/s10115-010-0375-z
复制
发表时间:
2012-01-01
影响因子:
2.7
通讯作者:
Czarnowski, Ireneusz
Czarnowski, Ireneusz
中科院分区:
计算机科学4区
文献类型:
--
作者:
Czarnowski, Ireneusz

文献摘要

被引文献

相似文献

监督机器学习中的实例选择,通常称为数据缩减,旨在决定应保留训练集中的哪些实例以供学习过程中进一步使用。实例选择可以提高学习模型的功能和泛化属性,缩短学习过程的时间,或者有助于扩展到大型数据源。本文提出了一种基于集群的实例选择方法,其学习过程由代理团队执行,并讨论了它的四种变体。基本假设是在训练数据分组后进行实例选择。为了验证所提出的方法并研究所使用的聚类方法对分类质量的影响,进行了计算实验。
Instance selection in the supervised machine learning, often referred to as the data reduction, aims at deciding which instances from the training set should be retained for further use during the learning process. Instance selection can result in increased capabilities and generalization properties of the learning model, shorter time of the learning process, or it can help in scaling up to large data sources. The paper proposes a cluster-based instance selection approach with the learning process executed by the team of agents and discusses its four variants. The basic assumption is that instance selection is carried out after the training data have been grouped into clusters. To validate the proposed approach and to investigate the influence of the clustering method used on the quality of the classification, the computational experiment has been carried out.