Safe semi-supervised learning based on weighted likelihood

Safe semi-supervised learning based on weighted likelihood
复制标题

DOI:
10.1016/j.neunet.2014.01.016
复制
发表时间:
2014-05-01
期刊:
影响因子:
7.8
通讯作者:
Takeuchi, Jun'ichi
Takeuchi, Jun'ichi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Kawakita, Masanori;Takeuchi, Jun'ichi

文献摘要

被引文献

相似文献

我们感兴趣的是开发一种在任何情况下都有效的安全的半监督学习。半监督学习假设除了n个有标记的数据外,还有n个未标记的数据可用。然而,几乎所有以前的半监督方法都需要额外的假设(不仅仅是未标记的数据)来改进监督学习。如果这些假设不满足,那么这些方法的表现可能比监督学习更差。Sokolovska, Cappe, and Yvon(2008)提出了一种基于加权似然方法的半监督方法。他们证明了在没有任何假设的情况下,这种方法的渐近性能永远不会比监督学习差(即,它是安全的)。他们的方法很有吸引力,因为它易于实现并且具有潜在的通用性。此外,它还与某种统计悖论密切相关。然而,Sokolovska et al.(2008)的方法假设了一个非常有限的情况,即分类、离散协变量、n'->∞和极大似然估计量。在本文中,我们通过修改权值来扩展他们的方法。我们证明,只要n
We are interested in developing a safe semi-supervised learning that works in any situation. Semi-supervised learning postulates that n' unlabeled data are available in addition to n labeled data. However, almost all of the previous semi-supervised methods require additional assumptions (not only unlabeled data) to make improvements on supervised learning. If such assumptions are not met, then the methods possibly perform worse than supervised learning. Sokolovska, Cappe, and Yvon (2008) proposed a semi-supervised method based on a weighted likelihood approach. They proved that this method asymptotically never performs worse than supervised learning (i.e., it is safe) without any assumption. Their method is attractive because it is easy to implement and is potentially general. Moreover, it is deeply related to a certain statistical paradox. However, the method of Sokolovska et al. (2008) assumes a very limited situation, i.e., classification, discrete covariates, n'-> infinity and a maximum likelihood estimator. In this paper, we extend their method by modifying the weight. We prove that our proposal is safe in a significantly wide range of situations as long as n