A Fast, Consistent Kernel Two-Sample Test

A Fast, Consistent Kernel Two-Sample Test
复制标题

DOI:
--
复制
发表时间:
2009-12
期刊:
--
影响因子:
--
通讯作者:
A. Gretton;K. Fukumizu;Zaïd Harchaoui;Bharath K. Sriperumbudur
A. Gretton;K. Fukumizu;Zaïd Harchaoui;Bharath K. Sriperumbudur
中科院分区:
其他
文献类型:
--
作者:
A. Gretton;K. Fukumizu;Zaïd Harchaoui;Bharath K. Sriperumbudur

文献摘要

被引文献

相似文献

最近提出了将概率分布嵌入到再生核希尔伯特空间(RKHS)中的核嵌入,它允许根据各自嵌入之间的距离来比较两个概率度量P和Q:对于足够丰富的RKHS,当且仅当P和Q重合时,该距离为零。在使用该距离作为检验两个样本是否来自不同分布的统计量时,在计算显著性阈值时出现了一个主要困难,因为经验统计量的零分布(其中P = Q)是χ2随机变量的无限加权和。先前的有限样本近似零分布包括使用自举回归,其产生一致的估计,但计算成本高;以及用检验统计量的低阶矩拟合参数模型,其在实践中可以很好地工作,但没有一致性或准确性保证。本工作的主要结果是一种新的估计的零分布,计算从Gram矩阵的特征谱的聚合样本从P和Q,并具有较低的计算成本比自助。这种估计的一致性的证明。零分布估计的性能进行了比较,与人工的例子,高维多变量数据和文本的引导和参数的方法。
A kernel embedding of probability distributions into reproducing kernel Hilbert spaces (RKHS) has recently been proposed, which allows the comparison of two probability measures P and Q based on the distance between their respective embeddings: for a sufficiently rich RKHS, this distance is zero if and only if P and Q coincide. In using this distance as a statistic for a test of whether two samples are from different distributions, a major difficulty arises in computing the significance threshold, since the empirical statistic has as its null distribution (where P = Q) an infinite weighted sum of χ2 random variables. Prior finite sample approximations to the null distribution include using bootstrap resampling, which yields a consistent estimate but is computationally costly; and fitting a parametric model with the low order moments of the test statistic, which can work well in practice but has no consistency or accuracy guarantees. The main result of the present work is a novel estimate of the null distribution, computed from the eigen-spectrum of the Gram matrix on the aggregate sample from P and Q, and having lower computational cost than the bootstrap. A proof of consistency of this estimate is provided. The performance of the null distribution estimate is compared with the bootstrap and parametric approaches on an artificial example, high dimensional multivariate data, and text.