A two-stage ensemble method for the detection of class-label noise

A two-stage ensemble method for the detection of class-label noise
复制标题

DOI:
10.1016/j.neucom.2017.11.012
复制
发表时间:
2018-01
期刊:
影响因子:
6
通讯作者:
Maryam Sabzevari;Gonzalo Martínez-Muñoz;A. Suárez
Maryam Sabzevari;Gonzalo Martínez-Muñoz;A. Suárez
中科院分区:
计算机科学2区
文献类型:
--
作者:
Maryam Sabzevari;Gonzalo Martínez-Muñoz;A. Suárez

文献摘要

被引文献

相似文献

Bootstrap 集成的属性(例如装袋或随机森林)用于检测和处理分类问题中的标签噪声。第一个观察结果是,子采样是一种正则化机制,可用于使自举集合对此类噪声更加鲁棒。此外,可以使用袋外数据来估计采样率的适当值。第二个观察结果是,集成分类器往往会在错误标记的实例中犯更多错误。因此,足够大比例的整体预测器出错的实例被标记为噪声。该阈值的合适值取决于问题,通过使用包装方法的交叉验证来确定。然后,可以过滤被识别为噪声的实例(即丢弃用于训练),或者通过更正其类标签来清理。最后,在这些清理后的训练数据上重新构建一个集成。对不同应用领域的分类问题进行的大量实验表明,即使存在高水平的类标签噪声,该过程也可以有效地构建准确的集成。
The properties of bootstrap ensembles, such as bagging or random forest, are utilized to detect and handle label noise in classification problems. The first observation is that subsampling is a regularization mechanism that can be used to render bootstrap ensembles more robust to this type of noise. Furthermore, appropriate values of the sampling rate can be estimated using out-of-bag data. A second observation is that the ensemble classifiers tend to make more errors in incorrectly labeled instances. Thus, instances for which a sufficiently large fraction of ensemble predictors err are marked as noisy. Suitable values of this threshold, which are problem dependent, are determined by cross-validation using a wrapper method. Instances identified as noisy can then be either filtered (i.e. discarded for training), or cleaned by correcting their class labels. Finally, an ensemble is built afresh on these cleansed training data. Extensive experiments in classification problems from different areas of application show that this procedure is effective to build accurate ensembles, even in the presence of high levels of class-label noise.