A Weighted Edge-Count Two-Sample Test for Multivariate and Object Data

A Weighted Edge-Count Two-Sample Test for Multivariate and Object Data
复制标题

DOI:
10.1080/01621459.2017.1307757
复制
发表时间:
2018-01-01
影响因子:
3.7
通讯作者:
Su, Yi
Su, Yi
中科院分区:
数学1区
文献类型:
--
作者:
Chen, Hao;Chen, Xu;Su, Yi

文献摘要

被引文献

相似文献

多变量数据和非欧盟数据数据的两样本测试广泛用于许多领域。参数测试主要限制为符合参数模型假设的某些类型的数据。在本文中,我们研究了一种非参数测试程序,该程序使用代表观测值相似性的图表。只要可以定义样本空间上的信息相似度度量,它就可以应用于任何数据类型。当两个样本尺寸不同时,基于相似图的经典测试会出现问题。我们通过将适当的权重应用于经典测试统计量的不同组件来解决问题。新的测试在模拟研究中表现出了可观的功率增长。其渐近排列的无效分布被得出并证明在有限样品下可很好地工作,从而促进了其在大型数据集中的应用。通过对网络数据的真实数据集进行分析来说明新测试。
Two-sample tests for multivariate data and non-Euclidean data are widely used in many fields. Parametric tests are mostly restrained to certain types of data that meets the assumptions of the parametric models. In this article, we study a nonparametric testing procedure that uses graphs representing the similarity among observations. It can be applied to any data types as long as an informative similarity measure on the sample space can be defined. The classic test based on a similarity graph has a problem when the two sample sizes are different. We solve the problem by applying appropriate weights to different components of the classic test statistic. The new test exhibits substantial power gains in simulation studies. Its asymptotic permutation null distribution is derived and shown to work well under finite samples, facilitating its application to large datasets. The new test is illustrated through an analysis on a real dataset of network data.