Class Prior Estimation from Positive and Unlabeled Data

Class Prior Estimation from Positive and Unlabeled Data
复制标题

DOI:
10.1587/transinf.e97.d.1358
复制
发表时间:
2014-05
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
M. C. D. Plessis;Masashi Sugiyama
M. C. D. Plessis;Masashi Sugiyama
中科院分区:
其他
文献类型:
--
作者:
M. C. D. Plessis;Masashi Sugiyama

文献摘要

被引文献

相似文献

我们考虑只使用阳性和未标记样本学习分类器的问题。在这种设置中,已知如果类先验可用,则可以成功地学习分类器。然而,在实践中,类先验是未知的,因此必须从数据中估计。在本文中,我们提出了一种新的方法来估计类先验的部分匹配的类条件密度的正类的输入密度。通过在皮尔逊散度方面执行这种部分匹配,我们通过下界最大化直接估计而无需密度估计,我们可以获得类先验的分析估计。我们进一步表明,现有的类先验估计方法也可以被解释为皮尔逊发散下进行部分匹配,但在一个间接的方式。我们的直接类先验估计方法的优越性说明了几个基准数据集。
We consider the problem of learning a classifier using only positive and unlabeled samples. In this setting, it is known that a classifier can be successfully learned if the class prior is available. However, in practice, the class prior is unknown and thus must be estimated from data. In this paper, we propose a new method to estimate the class prior by partially matching the class-conditional density of the positive class to the input density. By performing this partial matching in terms of the Pearson divergence, which we estimate directly without density estimation via lower-bound maximization, we can obtain an analytical estimator of the class prior. We further show that an existing class prior estimation method can also be interpreted as performing partial matching under the Pearson divergence, but in an indirect manner. The superiority of our direct class prior estimation method is illustrated on several benchmark datasets.