A new framework for distance and kernel-based metrics in high dimensions

A new framework for distance and kernel-based metrics in high dimensions
复制标题

DOI:
10.1214/21-ejs1889
复制
发表时间:
2021-01-01
影响因子:
1.1
通讯作者:
Zhang, Xianyang
Zhang, Xianyang
中科院分区:
数学3区
文献类型:
--
作者:
Chakraborty, Shubhadeep;Zhang, Xianyang

文献摘要

被引文献

相似文献

本文提出了新的度量来量化和测试(i)分布的平等性和(ii)两个高维随机向量之间的独立性。我们发现,通常的欧几里得距离的基础上的能量距离不能完全表征的均匀性的两个高维分布的意义上说,它只检测的平等的手段和协方差矩阵的痕迹在高维设置。我们提出了一类新的度量,它继承了理想的性能的能量距离和最大平均差异/(广义)距离协方差和希尔伯特-施密特独立性准则在低维设置,并能够检测的同质性/完全表征之间的独立性在高维设置的低维边缘分布。我们进一步提出了基于新度量的t检验来进行高维双样本检验/独立性检验,并研究了它们在高维低样本量(HDLSS)和高维中样本量(HDMSS)设置下的渐近行为。t检验的计算复杂度仅随维度线性增长,因此可扩展到非常高维的数据。我们通过模拟和真实的数据集证明了所提出的分布均匀性和独立性检验的上级功效行为。
The paper presents new metrics to quantify and test for (i) the equality of distributions and (ii) the independence between two high-dimensional random vectors. We show that the energy distance based on the usual Euclidean distance cannot completely characterize the homogeneity of two high-dimensional distributions in the sense that it only detects the equality of means and the traces of covariance matrices in the high-dimensional setup. We propose a new class of metrics which inherits the desirable properties of the energy distance and maximum mean discrepancy/(generalized) distance covariance and the Hilbert-Schmidt Independence Criterion in the low-dimensional setting and is capable of detecting the homogeneity of/completely characterizing independence between the low-dimensional marginal distributions in the high dimensional setup. We further propose t-tests based on the new metrics to perform high-dimensional two-sample testing/independence testing and study their asymptotic behavior under both high dimension low sample size (HDLSS) and high dimension medium sample size (HDMSS) setups. The computational complexity of the t-tests only grows linearly with the dimension and thus is scalable to very high dimensional data. We demonstrate the superior power behavior of the proposed tests for homogeneity of distributions and independence via both simulated and real datasets.