Semi-Supervised Classifications via Elastic and Robust Embedding

Semi-Supervised Classifications via Elastic and Robust Embedding
复制标题

DOI:
10.1609/aaai.v31i1.10946
复制
发表时间:
2017-02
期刊:
--
影响因子:
--
通讯作者:
Yun Liu;Yiming Guo;Hua Wang;F. Nie;Heng Huang
Yun Liu;Yiming Guo;Hua Wang;F. Nie;Heng Huang
中科院分区:
其他
文献类型:
--
作者:
Yun Liu;Yiming Guo;Hua Wang;F. Nie;Heng Huang

文献摘要

被引文献

相似文献

直导式半监督学习只能对训练数据中出现的未标注数据进行标签预测,而不能对训练集中未出现的测试数据进行标注预测。为了处理这种样本外问题,许多归纳方法提出了一个约束,使得预测的标签矩阵应该精确地等于线性模型。在实践中,这种约束可能过于严格,无法捕获数据的多种结构。在本文中,我们放松了这一刚性约束,并建议在预测的标签矩阵上使用弹性约束,以便更好地探索流形结构。此外,由于实际中未标记的数据往往非常丰富,并且通常存在一些异常值,因此我们使用非平方损失而不是传统的平方损失来学习稳健模型。导出的问题虽然是凸的,但含有如此多的非光滑项,这使得它的求解非常具有挑战性。在本文中,我们提出了一种有效的优化算法来解决更一般的问题,并在此基础上找到了派生问题的最优解。
Transductive semi-supervised learning can only predict labels for unlabeled data appearing in training data, and can not predict labels for testing data never appearing in training set. To handle this out-of-sample problem, many inductive methods make a constraint such that the predicted label matrix should be exactly equal to a linear model. In practice, this constraint might be too rigid to capture the manifold structure of data. In this paper, we relax this rigid constraint and propose to use an elastic constraint on the predicted label matrix such that the manifold structure can be better explored. Moreover, since unlabeled data are often very abundant in practice and usually there are some outliers, we use a non-squared loss instead of the traditional squared loss to learn a robust model. The derived problem, although is convex, has so many nonsmooth terms, which make it very challenging to solve. In the paper, we propose an efficient optimization algorithm to solve a more general problem, based on which we find the optimal solution to the derived problem.