Large data and zero noise limits of graph-based semi-supervised learning algorithms

Large data and zero noise limits of graph-based semi-supervised learning algorithms
复制标题

DOI:
10.1016/j.acha.2019.03.005
复制
发表时间:
2020-09-01
影响因子:
2.5
通讯作者:
Thorpe, Matthew
Thorpe, Matthew
中科院分区:
数学1区
文献类型:
--
作者:
Dunlop, Matthew M.;Slepcev, Dejan;Thorpe, Matthew

文献摘要

被引文献

相似文献

图拉普拉斯在大图极限下接近微分算子的缩放被用来发展对一些半监督学习算法的理解;重点介绍了概率算法、水平集和克里格方法。优化和贝叶斯方法都被考虑,基于从拉普拉斯仿射变换中发现的正则二次型,提高到可能的分数指数。确定了定义该二次型的参数的条件,在此条件下,可以找到优化和贝叶斯半监督学习问题的良好定义的极限连续体类似物,从而为大型图集中的算法设计提供了启发。使用最近引入的TLp度量,通过g收敛解决了优化公式的大图极限。本文还确定了贝叶斯公式的小标记噪声限制,并与已有的调和函数方法进行了对比。(C) 2019 Elsevier Inc.版权所有。
Scalings in which the graph Laplacian approaches a differential operator in the large graph limit are used to develop understanding of a number of algorithms for semi-supervised learning; in particular, the probit algorithm, level set and kriging methods. Both optimization and Bayesian approaches are considered, based around a regularizing quadratic form found from an affine transformation of the Laplacian, raised to a possibly fractional, exponent. Conditions on the parameters defining this quadratic form are identified under which well-defined limiting continuum analogues of the optimization and Bayesian semi-supervised learning problems may be found, thereby shedding light on the design of algorithms in the large graph setting. The large graph limits of the optimization formulations are tackled through G-convergence, using the recently introduced TLp metric. The small labeling noise limits of the Bayesian formulations are also identified, and contrasted with pre-existing harmonic function approaches to the problem. (C) 2019 Elsevier Inc. All rights reserved.