Learning Bayesian classifiers from positive and unlabeled examples

Learning Bayesian classifiers from positive and unlabeled examples
复制标题

DOI:
10.1016/j.patrec.2007.08.003
复制
发表时间:
2007-12-01
影响因子:
5.1
通讯作者:
Lozano, Jose A.
Lozano, Jose A.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Calvo, Boria;Larranaga, Pedro;Lozano, Jose A.

文献摘要

被引文献

相似文献

正无标记学习项是指在没有负样本的情况下的二元分类问题。当只有阳性和未标记的实例可用时,半监督分类算法不能直接应用,因此需要新的算法。这些积极的未标记的学习算法之一是积极的朴素贝叶斯(PNB),这是一个适应的朴素贝叶斯归纳算法,不需要负的情况。在这项工作中,我们提出了两种方法来增强这种算法。一方面,我们将PNB背后的概念更进一步,提出了一个在没有负面实例的情况下构建更复杂的贝叶斯分类器的过程。本文提出了一种新的算法(称为正树增广朴素贝叶斯,PTAN),以获得树增广朴素贝叶斯模型的积极的未标记的区域。另一方面,我们提出了一个新的贝叶斯方法来处理的先验概率的正类模型的不确定性,通过一个Beta分布。这种方法适用于PNB和PTAN,从而产生两个新的算法。在基于真实的和合成数据库的正无标记学习问题中,对这四种算法进行了经验比较。在这些比较中得到的结果表明,当预测变量不是条件独立的类,PNB的扩展到更复杂的网络增加了分类性能。他们还表明,我们的贝叶斯方法的先验概率的正类可以改善PNB和PTAN得到的结果。(c)2007 Elsevier B.V.保留所有权利。
The positive unlabeled learning term refers to the binary classification problem in the absence of negative examples. When only positive and unlabeled instances are available, semi-supervised classification algorithms cannot be directly applied, and thus new algorithms are required. One of these positive unlabeled learning algorithms is the positive naive Bayes (PNB), which is an adaptation of the naive Bayes induction algorithm that does not require negative instances. In this work we propose two ways of enhancing this algorithm. On one hand, we have taken the concept behind PNB one step further, proposing a procedure to build more complex Bayesian classifiers in the absence of negative instances. We present a new algorithm (named positive tree augmented naive Bayes, PTAN) to obtain tree augmented naive Bayes models in the positive unlabeled domain. On the other hand, we propose a new Bayesian approach to deal with the a priori probability of the positive class that models the uncertainty over this parameter by means of a Beta distribution. This approach is applied to both PNB and PTAN, resulting in two new algorithms. The four algorithms are empirically compared in positive unlabeled learning problems based on real and synthetic databases. The results obtained in these comparisons suggest that, when the predicting variables are not conditionally independent given the class, the extension of PNB to more complex networks increases the classification performance. They also show that our Bayesian approach to the a priori probability of the positive class can improve the results obtained by PNB and PTAN. (c) 2007 Elsevier B.V. All rights reserved.