False positive rate control for positive unlabeled learning

False positive rate control for positive unlabeled learning
复制标题

阳性无标记学习的误报率控制

DOI:
10.1016/j.neucom.2019.08.001
复制
发表时间:
2019
期刊:
影响因子:
6
通讯作者:
Wang Jun
Wang Jun
中科院分区:
计算机科学2区
文献类型:
--
作者:
Kong Shuchen;Shen Weiwei;Zheng Yingbin;Zhang Ao;Pu Jian;Wang Jun

文献摘要

被引文献

相似文献

具有误检率控制的学习分类器在过去几年的应用中引起了广泛的关注。虽然已经开发了各种监督算法来获得低的假阳性率,但它们通常要求数据中同时存在正样本和负样本。然而,在正向无标记(PU)学习中研究的情景在实践中更为普遍。也就是说,在开始时,大多数数据可能不具有已知标签,并且具有已知标签的数据可能仅代表一种类型的样本。为了应对这一挑战,本文提出了一种新的具有误检率控制的阳性无标记学习分类器。特别地,我们首先证明了在这种情况下,使用通常采用的凸代换损失函数,例如铰链损失函数,会产生对错误正确率的冗余惩罚。然后,我们提出了非凸斜坡损失代理函数可以克服这一障碍,并证明了凹凸法可以解决相关的非凸优化问题。最后,通过在多个数据集上的大量实验,验证了该方法的有效性。
Learning classifiers with false positive rate control have drawn intensive attention in applications over past years. While various supervised algorithms have been developed for obtaining low false positive rates, they commonly require the coexistence of both positive and negative samples in data. However, the scenario studied in positive unlabeled (PU) learning is more pervasive in practice. Namely, at inception, most of the data may not have known labels, and the data with known labels may only represent one type of samples. To tackle this challenge, in this paper we propose a new positive unlabeled learning classifier with false positive rate control. In particular, we first prove that in this context employing oft-adopted convex surrogate loss functions, such as the hinge loss function, begets a redundant penalty for false positive rates. Then, we present that the non-convex ramp loss surrogate function can overcome this barrier and show a concave-convex procedure can solve the associated non-convex optimization problem. Finally, we demonstrate the effectiveness of the proposed method through extensive experiments on multiple datasets.