Randomized incomplete $U$-statistics in high dimensions

Randomized incomplete $U$-statistics in high dimensions
复制标题

DOI:
10.1214/18-aos1773
复制
发表时间:
2017-12
期刊:
The Annals of Statistics
影响因子:
--
通讯作者:
Xiaohui Chen;Kengo Kato
Xiaohui Chen;Kengo Kato
中科院分区:
其他
文献类型:
--
作者:
Xiaohui Chen;Kengo Kato

文献摘要

被引文献

相似文献

本文研究了高维$U$-统计量的平均向量的推断。在大数据时代,U统计量的维数d和观测值的样本量n都趋于较大,U统计量的计算要求过高。依赖于数据的推理过程,如U统计的经验引导,在计算上更加昂贵。为了克服这样的计算瓶颈,通过采样较少的U统计项获得的不完全U统计是一个有吸引力的选择。本文引入具有稀疏权值的随机化不完全$U$-统计量,其计算代价与$U$-统计量的阶数无关。我们推导了高维随机不完全U统计量的非渐近高斯近似误差界,即在维数d可能远远大于样本大小n的情况下,对于非简并核和简并核。此外,我们提出了不完全$U$-统计量的通用自举方法,这些方法的计算要求比现有的自举方法低得多,并建立了所提出的自举方法的有限样本有效性。我们的方法说明了应用于非参数检验的两两独立性的高维随机向量在较弱的假设比那些出现在文献。
This paper studies inference for the mean vector of a high-dimensional $U$-statistic. In the era of Big Data, the dimension $d$ of the $U$-statistic and the sample size $n$ of the observations tend to be both large, and the computation of the $U$-statistic is prohibitively demanding. Data-dependent inferential procedures such as the empirical bootstrap for $U$-statistics is even more computationally expensive. To overcome such computational bottleneck, incomplete $U$-statistics obtained by sampling fewer terms of the $U$-statistic are attractive alternatives. In this paper, we introduce randomized incomplete $U$-statistics with sparse weights whose computational cost can be made independent of the order of the $U$-statistic. We derive non-asymptotic Gaussian approximation error bounds for the randomized incomplete $U$-statistics in high dimensions, namely in cases where the dimension $d$ is possibly much larger than the sample size $n$, for both non-degenerate and degenerate kernels. In addition, we propose generic bootstrap methods for the incomplete $U$-statistics that are computationally much less-demanding than existing bootstrap methods, and establish finite sample validity of the proposed bootstrap methods. Our methods are illustrated on the application to nonparametric testing for the pairwise independence of a high-dimensional random vector under weaker assumptions than those appearing in the literature.