Analysis of new techniques to obtain quality training sets
Analysis of new techniques to obtain quality training sets
复制标题
DOI:
10.1016/s0167-8655(02)00225-8
复制
发表时间:
2003-04
期刊:
影响因子:
--
通讯作者:
J. S. Sánchez;R. Barandela;A. I. Marqués;R. Alejo;J. Badenas
中科院分区:
文献类型:
--
作者:
J. S. Sánchez;R. Barandela;A. I. Marqués;R. Alejo;J. Badenas
This paper presents new algorithms to identify and eliminate mislabelled, noisy and atypical training samples for supervised learning and more specifically, for nearest neighbour classification. The main goal of these approaches is to enhance the classification accuracy by improving the quality of the training data. Several experiments with synthetic and real data sets are carried out in order to illustrate the behaviour of the schemes proposed here and compare their performance with that of other traditional techniques. It is also analysed the ability of these new algorithms to “reduce” the possible overlapping among regions of different classes.