A Normality Test for High-dimensional Data Based on the Nearest Neighbor Approach

A Normality Test for High-dimensional Data Based on the Nearest Neighbor Approach
复制标题

DOI:
10.1080/01621459.2021.1953507
复制
发表时间:
2019-04
影响因子:
3.7
通讯作者:
Hao Chen;Yin Xia
Hao Chen;Yin Xia
中科院分区:
数学1区
文献类型:
--
作者:
Hao Chen;Yin Xia

文献摘要

被引文献

相似文献

摘要 许多高维数据的统计方法都假设总体是正态的。尽管已经提出了一些多元正态性检验,但据我们所知,当维度大于观测值数量时,它们都无法正确控制 I 类错误。在这项工作中,我们提出了一种使用最近邻信息的新型非参数测试。该方法保证了高维设置下的渐近I型误差控制。仿真研究验证了当尺寸随着样本大小而增长时所提出的测试的经验尺寸性能,同时与替代方法相比,新测试表现出优越的功率性能。我们还通过高维分类和聚类文献中两个常用的数据集来说明我们的方法,其中偏离正态性假设可能会导致无效的结论。
Abstract Many statistical methodologies for high-dimensional data assume the population is normal. Although a few multivariate normality tests have been proposed, to the best of our knowledge, none of them can properly control the Type I error when the dimension is larger than the number of observations. In this work, we propose a novel nonparametric test that uses the nearest neighbor information. The proposed method guarantees the asymptotic Type I error control under the high-dimensional setting. Simulation studies verify the empirical size performance of the proposed test when the dimension grows with the sample size and at the same time exhibit a superior power performance of the new test compared with alternative methods. We also illustrate our approach through two popularly used datasets in high-dimensional classification and clustering literatures where deviation from the normality assumption may lead to invalid conclusions.