Kernel Based Detection of Mislabeled Training Examples

Kernel Based Detection of Mislabeled Training Examples
复制标题

DOI:
10.1137/1.9781611972771.28
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
Hamed Valizadegan;P. Tan
Hamed Valizadegan;P. Tan
中科院分区:
其他
文献类型:
--
作者:
Hamed Valizadegan;P. Tan

文献摘要

被引文献

相似文献

识别错误标记的训练样本的问题已经在几项研究中进行了研究,开发了各种方法来编辑训练数据以获得更好的分类器。这些方法中的许多方法涉及将单个或一组分类器应用于训练集,并基于它们相对于分类器输出的一致性来过滤错误标记的样本。在本研究中,我们将错标检测问题描述为一个最优化问题,并引入了一种基于核的方法来过滤错标样本。使用UCI数据仓库中的各种数据集进行的实验结果表明,与现有的基于最近邻和集成的过滤方案相比,我们提出的方法是有效的。
The problem of identifying mislabeled training examples has been examined in several studies, with a variety of approaches developed for editing the training data to obtain better classifiers. Many of these approaches involve applying an individual or an ensemble of classifiers to the training set and filtering the mislabeled examples based on their consistency with respect to the classifier’s outputs. In this study, we formulate mislabeled detection as an optimization problem and introduce a kernel-based approach for filtering the mislabeled examples. Experimental results using a variety of data sets from the UCI data repository demonstrate the effectiveness of our proposed method, compared to existing nearest-neighbor and ensemble-based filtering schemes.