Positive-unlabeled classification under class-prior shift: a prior-invariant approach based on density ratio estimation

Positive-unlabeled classification under class-prior shift: a prior-invariant approach based on density ratio estimation
复制标题

DOI:
10.1007/s10994-022-06190-z
复制
发表时间:
2021-07
期刊:
影响因子:
7.5
通讯作者:
Shōta Nakajima;Masashi Sugiyama
Shōta Nakajima;Masashi Sugiyama
中科院分区:
计算机科学3区
文献类型:
--
作者:
Shōta Nakajima;Masashi Sugiyama

文献摘要

相似文献

从正的和未标记的(PU)数据中学习是各种应用中的一个重要问题。最近的PU分类方法大多假设训练未标记数据集中的类先验(阳性样本的比例)与测试数据中的类先验(阳性样本的比例)相同,这在许多实际情况下并不成立。此外,我们通常不知道训练和测试数据的类先验,因此我们不知道如何在没有它们的情况下训练分类器。为了解决这些问题,我们提出了一种新的PU分类方法的基础上密度比估计。我们所提出的方法的一个显着的优点是,它不需要在训练阶段的类先验;类先验移位只在测试阶段。我们从理论上证明我们提出的方法,并通过实验证明其有效性。
Learning from positive and unlabeled (PU) data is an important problem in various applications. Most of the recent approaches for PU classification assume that the class-prior (the ratio of positive samples) in the training unlabeled dataset is identical to that of the test data, which does not hold in many practical cases. In addition, we usually do not know the class-priors of the training and test data, thus we have no clue on how to train a classifier without them. To address these problems, we propose a novel PU classification method based on density ratio estimation. A notable advantage of our proposed method is that it does not require the class-priors in the training phase; class-prior shift is incorporated only in the test phase. We theoretically justify our proposed method and experimentally demonstrate its effectiveness.