Bayesian Kernel Two-Sample Testing

Bayesian Kernel Two-Sample Testing
复制标题

DOI:
10.1080/10618600.2022.2067547
复制
发表时间:
2020-02
影响因子:
2.4
通讯作者:
Qinyi Zhang;S. Filippi;S. Flaxman;D. Sejdinovic
Qinyi Zhang;S. Filippi;S. Flaxman;D. Sejdinovic
中科院分区:
数学2区
文献类型:
--
作者:
Qinyi Zhang;S. Filippi;S. Flaxman;D. Sejdinovic

文献摘要

相似文献

摘要在现代数据分析中,随机变量间差异的非参数度量尤为重要。这一主题在频率主义文献中得到了很好的研究,而在贝叶斯环境下的发展有限,其应用通常局限于单变量情况。这里,我们利用Flaxman等人建立的框架,提出了一种贝叶斯核两样本测试方法,该方法基于对再生核Hilbert空间中核均值嵌入的差异进行建模。核方法的使用使其能够应用于多变量欧氏空间以外的一般区域中的随机变量。所提出的过程导致了允许自动选择与手头问题相关的核参数的后验推断方案。通过一系列合成实验和两个真实数据实验(即从高维数据测试网络异构性和六元单环构象比较),我们展示了该方法的优势。这篇文章的补充材料可以在网上找到。
Abstract In modern data analysis, nonparametric measures of discrepancies between random variables are particularly important. The subject is well-studied in the frequentist literature, while the development in the Bayesian setting is limited where applications are often restricted to univariate cases. Here, we propose a Bayesian kernel two-sample testing procedure based on modeling the difference between kernel mean embeddings in the reproducing kernel Hilbert space using the framework established by Flaxman et al. The use of kernel methods enables its application to random variables in generic domains beyond the multivariate Euclidean spaces. The proposed procedure results in a posterior inference scheme that allows an automatic selection of the kernel parameters relevant to the problem at hand. In a series of synthetic experiments and two real data experiments (i.e., testing network heterogeneity from high-dimensional data and six-membered monocyclic ring conformation comparison), we illustrate the advantages of our approach. Supplementary materials for this article are available online.