Bayesian Learning of Kernel Embeddings

Bayesian Learning of Kernel Embeddings
复制标题

DOI:
--
复制
发表时间:
2016-03
期刊:
arXiv: Machine Learning
影响因子:
--
通讯作者:
S. Flaxman;D. Sejdinovic;J. Cunningham;S. Filippi
S. Flaxman;D. Sejdinovic;J. Cunningham;S. Filippi
中科院分区:
其他
文献类型:
--
作者:
S. Flaxman;D. Sejdinovic;J. Cunningham;S. Filippi

文献摘要

相似文献

核方法是机器学习的支柱之一,但核学习的问题仍然具有挑战性,只有一些启发式方法和很少的理论。这在基于概率度量的核均值嵌入估计的方法中特别重要。对于特征核(包括最常用的核),核均值嵌入唯一地确定了其概率度量,因此可以用来设计强大的统计测试框架,其中包括非参数双样本和独立性测试。然而,实际上,这些测试的性能对内核及其长度参数的选择非常敏感。为了解决这个核心问题,我们提出了一种新的核均值嵌入概率模型,即贝叶斯核嵌入模型,该模型将包含均值嵌入的再现核希尔伯特空间先验的高斯过程与共轭似然函数相结合,从而产生均值嵌入的封闭形式后验。我们模型的后验均值与最近提出的核均值嵌入的收缩估计器密切相关,而后验不确定性是一个新的、有趣的特征,具有各种可能的应用。至关重要的是,为了核学习的目的,我们的模型在给定核超参数的情况下给出了观察数据的简单、封闭形式的边际伪似然。这种边际伪似然可以被优化以告知超参数选择,也可以使用完全贝叶斯推理。
Kernel methods are one of the mainstays of machine learning, but the problem of kernel learning remains challenging, with only a few heuristics and very little theory. This is of particular importance in methods based on estimation of kernel mean embeddings of probability measures. For characteristic kernels, which include most commonly used ones, the kernel mean embedding uniquely determines its probability measure, so it can be used to design a powerful statistical testing framework, which includes nonparametric two-sample and independence tests. In practice, however, the performance of these tests can be very sensitive to the choice of kernel and its lengthscale parameters. To address this central issue, we propose a new probabilistic model for kernel mean embeddings, the Bayesian Kernel Embedding model, combining a Gaussian process prior over the Reproducing Kernel Hilbert Space containing the mean embedding with a conjugate likelihood function, thus yielding a closed form posterior over the mean embedding. The posterior mean of our model is closely related to recently proposed shrinkage estimators for kernel mean embeddings, while the posterior uncertainty is a new, interesting feature with various possible applications. Critically for the purposes of kernel learning, our model gives a simple, closed form marginal pseudolikelihood of the observed data given the kernel hyperparameters. This marginal pseudolikelihood can either be optimized to inform the hyperparameter choice or fully Bayesian inference can be used.