Mixture Proportion Estimation and PU Learning: A Modern Approach

Mixture Proportion Estimation and PU Learning: A Modern Approach
复制标题

DOI:
--
复制
发表时间:
2021-11
期刊:
ArXiv
影响因子:
--
通讯作者:
S. Garg;Yifan Wu;Alexander J. Smola;Sivaraman Balakrishnan;Zachary Chase Lipton
S. Garg;Yifan Wu;Alexander J. Smola;Sivaraman Balakrishnan;Zachary Chase Lipton
中科院分区:
其他
文献类型:
--
作者:
S. Garg;Yifan Wu;Alexander J. Smola;Sivaraman Balakrishnan;Zachary Chase Lipton

文献摘要

被引文献

相似文献

只给出肯定的例子和未标记的例子(来自肯定类和否定类),我们仍然希望估计出一个准确的肯定与否定分类器。从形式上讲,这个任务被分解为两个子任务:(i)混合比例估计(MPE)-确定未标记数据中阳性样本的比例;以及(ii)PU学习-给定这样的估计,学习所需的阳性与阴性分类器。不幸的是,这两个问题的经典方法在高维环境中崩溃。与此同时,最近提出的算法缺乏理论上的一致性,并且不稳定地依赖于超参数调整。在本文中,我们提出了两种简单的技术:最佳箱估计(BBE)(用于MPE);和条件值忽略风险(CVIR),这是PU学习的一个简单目标。这两种方法在经验上都主导了以前的方法,对于BBE,我们建立了正式的保证,只要我们可以训练模型,就可以干净地分离出一小部分正面示例。我们的最终算法(TED)$^n$,在两个过程之间交替,显着改善了我们的混合比例估计器和分类器
Given only positive examples and unlabeled examples (from both positive and negative classes), we might hope nevertheless to estimate an accurate positive-versus-negative classifier. Formally, this task is broken down into two subtasks: (i) Mixture Proportion Estimation (MPE) -- determining the fraction of positive examples in the unlabeled data; and (ii) PU-learning -- given such an estimate, learning the desired positive-versus-negative classifier. Unfortunately, classical methods for both problems break down in high-dimensional settings. Meanwhile, recently proposed heuristics lack theoretical coherence and depend precariously on hyperparameter tuning. In this paper, we propose two simple techniques: Best Bin Estimation (BBE) (for MPE); and Conditional Value Ignoring Risk (CVIR), a simple objective for PU-learning. Both methods dominate previous approaches empirically, and for BBE, we establish formal guarantees that hold whenever we can train a model to cleanly separate out a small subset of positive examples. Our final algorithm (TED)$^n$, alternates between the two procedures, significantly improving both our mixture proportion estimator and classifier