Uncertainty Quantification in Graph-Based Classification of High Dimensional Data

Uncertainty Quantification in Graph-Based Classification of High Dimensional Data
复制标题

DOI:
10.1137/17m1134214
复制
发表时间:
2018-01-01
影响因子:
2
通讯作者:
Zygalakis, Konstantinos C.
Zygalakis, Konstantinos C.
中科院分区:
工程技术3区
文献类型:
--
作者:
Bertozzi, Andrea L.;Luo, Xiyang;Zygalakis, Konstantinos C.

文献摘要

被引文献

相似文献

高维数据的分类有着广泛的应用。在许多这些应用中,为所得分类配备不确定性度量可能与分类本身一样重要。在本文中,我们介绍,开发算法,并调查各种贝叶斯模型的二进制分类任务的属性;通过分类标签上的后验分布,这些方法自动给出不确定性的措施。这些方法都是基于半监督学习的图形式。我们提供了一个统一的框架,汇集了各种方法,已在不同的社区内的数学科学。我们研究了概率单位分类[C. K.威廉姆斯和C. E. Rasmussen,“Gaussian Processes for Regression”,Advances in Neural Information Processing Systems 8,MIT Press,1996,pp. 514-520]在基于图的背景下,推广了贝叶斯反问题的水平集方法[M. A. Iglesias,Y. Lu和A. M.斯图尔特,接口自由绑定。18(2016),pp. 181-217]的分类设置,并推广Ginzburg-Landau优化为基础的分类器[A. L. Bertozzi和A. Flenner,多尺度模型。同时,10(2012),pp. 1090-1118],[Y.货车Gennip和A. L. Bertozzi,Adv.微分方程,17(2012),pp. 1115-1180]到贝叶斯设置。我们还表明,概率单位和水平集方法是[X]中引入的调和函数方法的自然松弛。Zhu等人,“Semi-supervised Learning Using Gaussian Fields and Harmonic Functions,”in ICML,Vol. 3,2003,pp. 912-919]。我们介绍了有效的数值方法,适合于大型数据集,MCMC为基础的采样和基于梯度的MAP估计。通过数值实验,我们研究了我们的模型的分类精度和不确定性量化;这些实验展示了一套通常用于评估基于图的半监督学习算法的数据集。
Classification of high dimensional data finds wide-ranging applications. In many of these applications equipping the resulting classification with a measure of uncertainty may be as important as the classification itself. In this paper we introduce, develop algorithms for, and investigate the properties of a variety of Bayesian models for the task of binary classification; via the posterior distribution on the classification labels, these methods automatically give measures of uncertainty. The methods are all based on the graph formulation of semisupervised learning. We provide a unified framework which brings together a variety of methods that have been introduced in different communities within the mathematical sciences. We study probit classification [C. K. Williams and C. E. Rasmussen, "Gaussian Processes for Regression," in Advances in Neural Information Processing Systems 8, MIT Press, 1996, pp. 514-520] in the graph-based setting, generalize the level-set method for Bayesian inverse problems [M. A. Iglesias, Y. Lu, and A. M. Stuart, Interfaces Free Bound., 18 (2016), pp. 181-217] to the classification setting, and generalize the Ginzburg-Landau optimization-based classifier [A. L. Bertozzi and A. Flenner, Multiscale Model. Simul., 10 (2012), pp. 1090-1118], [Y. Van Gennip and A. L. Bertozzi, Adv. Differential Equations, 17 (2012), pp. 1115-1180] to a Bayesian setting. We also show that the probit and level-set approaches are natural relaxations of the harmonic function approach introduced in [X. Zhu et al., "Semi-supervised Learning Using Gaussian Fields and Harmonic Functions," in ICML, Vol. 3, 2003, pp. 912-919]. We introduce efficient numerical methods, suited to large datasets, for both MCMC-based sampling and gradient-based MAP estimation. Through numerical experiments we study classification accuracy and uncertainty quantification for our models; these experiments showcase a suite of datasets commonly used to evaluate graph-based semisupervised learning algorithms.