Class-prior estimation for learning from positive and unlabeled data

Class-prior estimation for learning from positive and unlabeled data
复制标题

DOI:
10.1007/s10994-016-5604-6
复制
发表时间:
2017-04-01
期刊:
影响因子:
7.5
通讯作者:
Sugiyama, Masashi
Sugiyama, Masashi
中科院分区:
计算机科学3区
文献类型:
--
作者:
du Plessis, Marthinus C.;Niu, Gang;Sugiyama, Masashi

文献摘要

被引文献

相似文献

我们考虑在未标记数据集中估计类先验的问题。在额外的标记数据集可用的假设下,可以通过将类数据分布的混合拟合到未标记数据分布来估计类先验。然而,在实践中,这种额外的标记数据集通常不可用。在本文中,我们表明,与额外的样本只来自积极的类,类先验的未标记的数据集可以正确估计。我们的核心思想是使用适当的惩罚分歧模型拟合,以消除由于缺乏负样本所造成的错误。我们进一步表明,使用惩罚距离给出了一个计算效率高的算法与解析解。从理论上分析了算法的一致性、稳定性和估计误差。最后,我们通过实验证明了该方法的实用性。
We consider the problem of estimating the class prior in an unlabeled dataset. Under the assumption that an additional labeled dataset is available, the class prior can be estimated by fitting a mixture of class-wise data distributions to the unlabeled data distribution. However, in practice, such an additional labeled dataset is often not available. In this paper, we show that, with additional samples coming only from the positive class, the class prior of the unlabeled dataset can be estimated correctly. Our key idea is to use properly penalized divergences for model fitting to cancel the error caused by the absence of negative samples. We further show that the use of the penalized -distance gives a computationally efficient algorithm with an analytic solution. The consistency, stability, and estimation error are theoretically analyzed. Finally, we experimentally demonstrate the usefulness of the proposed method.