Analysis of new techniques to obtain quality training sets

Analysis of new techniques to obtain quality training sets
复制标题

DOI:
10.1016/s0167-8655(02)00225-8
复制
发表时间:
2003-04
期刊:
Pattern Recognit. Lett.
影响因子:
--
通讯作者:
J. S. Sánchez;R. Barandela;A. I. Marqués;R. Alejo;J. Badenas
J. S. Sánchez;R. Barandela;A. I. Marqués;R. Alejo;J. Badenas
中科院分区:
其他
文献类型:
--
作者:
J. S. Sánchez;R. Barandela;A. I. Marqués;R. Alejo;J. Badenas

文献摘要

被引文献

相似文献

本文提出了新的算法来识别和消除错误标记、噪声和非典型训练样本,用于监督学习,更具体地说,用于最近邻分类。这些方法的主要目标是通过提高训练数据的质量来提高分类精度。使用合成数据集和真实数据集进行了多次实验,以说明此处提出的方案的行为并将其性能与其他传统技术的性能进行比较。还分析了这些新算法“减少”不同类别区域之间可能重叠的能力。
This paper presents new algorithms to identify and eliminate mislabelled, noisy and atypical training samples for supervised learning and more specifically, for nearest neighbour classification. The main goal of these approaches is to enhance the classification accuracy by improving the quality of the training data. Several experiments with synthetic and real data sets are carried out in order to illustrate the behaviour of the schemes proposed here and compare their performance with that of other traditional techniques. It is also analysed the ability of these new algorithms to “reduce” the possible overlapping among regions of different classes.