Semi-Supervised Classification Based on Classification from Positive and Unlabeled Data

Semi-Supervised Classification Based on Classification from Positive and Unlabeled Data
复制标题

DOI:
--
复制
发表时间:
2016-05
期刊:
--
影响因子:
--
通讯作者:
Tomoya Sakai;M. C. D. Plessis;Gang Niu;Masashi Sugiyama
Tomoya Sakai;M. C. D. Plessis;Gang Niu;Masashi Sugiyama
中科院分区:
其他
文献类型:
--
作者:
Tomoya Sakai;M. C. D. Plessis;Gang Niu;Masashi Sugiyama

文献摘要

被引文献

相似文献

迄今为止开发的大多数半监督分类方法在特定的分布假设(如聚类假设)下使用未标记数据进行正则化。相比之下,最近开发的阳性和未标记数据分类方法(PU分类)使用未标记数据进行风险评估,即,直接从未标记的数据中提取标记信息。在本文中,我们扩展PU分类,也将负面数据,并提出了一种新的半监督分类方法。我们建立了我们的新方法的泛化误差界,并表明界相对于未标记数据的数量减少,而无需现有半监督分类方法所需的分布假设。通过实验,我们证明了所提出的方法的实用性。
Most of the semi-supervised classification methods developed so far use unlabeled data for regularization purposes under particular distributional assumptions such as the cluster assumption. In contrast, recently developed methods of classification from positive and unlabeled data (PU classification) use unlabeled data for risk evaluation, i.e., label information is directly extracted from unlabeled data. In this paper, we extend PU classification to also incorporate negative data and propose a novel semi-supervised classification approach. We establish generalization error bounds for our novel methods and show that the bounds decrease with respect to the number of unlabeled data without the distributional assumptions that are required in existing semi-supervised classification methods. Through experiments, we demonstrate the usefulness of the proposed methods.