Semisupervised Classification With Cluster Regularization

Semisupervised Classification With Cluster Regularization
复制标题

DOI:
10.1109/tnnls.2012.2214488
复制
发表时间:
2012-10
影响因子:
10.4
通讯作者:
Rodrigo G. F. Soares;Huanhuan Chen;X. Yao
Rodrigo G. F. Soares;Huanhuan Chen;X. Yao
中科院分区:
计算机科学1区
文献类型:
--
作者:
Rodrigo G. F. Soares;Huanhuan Chen;X. Yao

文献摘要

被引文献

相似文献

半监督分类(SSC)从廉价的未标记数据和标记数据中学习,以预测测试实例的标签。为了利用来自未标记数据的信息,在真实的类结构和数据分布之间应该有一个假设的关系。一个假设是,聚集在一起的数据点可能具有相同的类标签。在本文中,我们提出了一种新的算法,即,基于聚类的正则化(clusterReg)的SSC,它采取分区的聚类算法作为一个正则化项的损失函数的SSC分类。QuarterReg根据聚类结构和有限的标记数据进行预测。实验结果表明,该算法对实际问题具有较好的泛化能力。当数据遵循这种聚类假设时,它的性能非常好。即使这些聚类有误导性的重叠,它仍然优于其他最先进的算法。
Semisupervised classification (SSC) learns, from cheap unlabeled data and labeled data, to predict the labels of test instances. In order to make use of the information from unlabeled data, there should be an assumed relationship between the true class structure and the data distribution. One assumption is that data points clustered together are likely to have the same class label. In this paper, we propose a new algorithm, namely, cluster-based regularization (ClusterReg) for SSC, that takes the partition given by a clustering algorithm as a regularization term in the loss function of an SSC classifier. ClusterReg makes predictions according to the cluster structure together with limited labeled data. The experiments confirmed that ClusterReg has a good generalization ability for real-world problems. Its performance is excellent when data follows this cluster assumption. Even when these clusters have misleading overlaps, it still outperforms other state-of-the-art algorithms.