Classification with unknown class conditional label noise on non-compact feature spaces

Classification with unknown class conditional label noise on non-compact feature spaces
复制标题

非紧特征空间上未知类条件标签噪声的分类

DOI:
--
复制
发表时间:
2019
期刊:
Annual Conference Computational Learning Theory
影响因子:
--
通讯作者:
A. Kabán
A. Kabán
中科院分区:
--
文献类型:
--
作者:
Henry W. J. Reeve;A. Kabán

文献摘要

被引文献

相似文献

我们研究了存在未知类条件标签噪声的分类问题,其中学习者观察到的标签已被损坏,具有一些未知的类相关概率。为了获得有限的采样率,以前的方法与未知的类条件标签噪声的分类要求回归函数是接近其极值集的大措施。我们将在非紧度量空间中考虑这个问题,其中回归函数不必达到其极值。 在这种情况下,我们确定最小最大最优学习率(直到对数因子)。速率显示有趣的阈值行为:当回归函数以足够的速率接近其极值时,最佳学习速率与在无标签噪声设置中获得的学习速率具有相同的顺序。如果回归函数逐渐接近其极值,则分类性能必然会降低。此外,我们提出了一种自适应算法,达到这些速率没有先验知识的分布参数或局部密度。这是第一次识别出在标签噪声设置中可以实现有限采样率的情况,但它们不同于没有标签噪声的最佳速率。
We investigate the problem of classification in the presence of unknown class-conditional label noise in which the labels observed by the learner have been corrupted with some unknown class dependent probability. In order to obtain finite sample rates, previous approaches to classification with unknown class-conditional label noise have required that the regression function is close to its extrema on sets of large measure. We shall consider this problem in the setting of non-compact metric spaces, where the regression function need not attain its extrema. In this setting we determine the minimax optimal learning rates (up to logarithmic factors). The rate displays interesting threshold behaviour: When the regression function approaches its extrema at a sufficient rate, the optimal learning rates are of the same order as those obtained in the label-noise free setting. If the regression function approaches its extrema more gradually then classification performance necessarily degrades. In addition, we present an adaptive algorithm which attains these rates without prior knowledge of either the distributional parameters or the local density. This identifies for the first time a scenario in which finite sample rates are achievable in the label noise setting, but they differ from the optimal rates without label noise.