Laplacian Support Vector Machines Trained in the Primal

Laplacian Support Vector Machines Trained in the Primal
复制标题

DOI:
10.5555/1953048.2021038
复制
发表时间:
2009-09
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
S. Melacci;M. Belkin
S. Melacci;M. Belkin
中科院分区:
其他
文献类型:
--
作者:
S. Melacci;M. Belkin

文献摘要

被引文献

相似文献

在过去的几年里,由于未标记数据的日益普遍,机器学习社区已经花费了大量的精力来开发更好的理解和提高利用未标记数据的分类器的质量。在流形正则化方法之后,拉普拉斯支持向量机(LapSVM)在半监督分类中表现出了最先进的性能。在本文中,我们提出了两种策略来解决原始的LapSVM问题,以克服一些问题的原始对偶制定。特别是,训练LapSVM的原始可以有效地执行与预处理共轭梯度。我们通过使用基于未标记数据或标记验证示例(如果可用)的预测的早期停止策略来加速训练。这使得算法能够快速计算近似解,其分类精度与最佳解大致相同,从而大大减少了训练时间。训练算法的计算复杂度从O(n3)降低到O(kn2),其中n是标记和未标记示例的组合数量,并且k根据经验评估为显著小于n。由于其简单性,在原始中训练LapSVM可以作为原始LapSVM公式的额外增强的起点,例如处理大型数据集的起点。我们提出了一个广泛的实验评估真实的世界的数据显示所提出的方法的好处。
In the last few years, due to the growing ubiquity of unlabeled data, much effort has been spent by the machine learning community to develop better understanding and improve the quality of classifiers exploiting unlabeled data. Following the manifold regularization approach, Laplacian Support Vector Machines (LapSVMs) have shown the state of the art performance in semi-supervised classification. In this paper we present two strategies to solve the primal LapSVM problem, in order to overcome some issues of the original dual formulation. In particular, training a LapSVM in the primal can be efficiently performed with preconditioned conjugate gradient. We speed up training by using an early stopping strategy based on the prediction on unlabeled data or, if available, on labeled validation examples. This allows the algorithm to quickly compute approximate solutions with roughly the same classification accuracy as the optimal ones, considerably reducing the training time. The computational complexity of the training algorithm is reduced from O(n3) to O(kn2), where n is the combined number of labeled and unlabeled examples and k is empirically evaluated to be significantly smaller than n. Due to its simplicity, training LapSVM in the primal can be the starting point for additional enhancements of the original LapSVM formulation, such as those for dealing with large data sets. We present an extensive experimental evaluation on real world data showing the benefits of the proposed approach.