SIMPLE: Statistical inference on membership profiles in large networks

SIMPLE: Statistical inference on membership profiles in large networks
复制标题

简单:对大型网络中的成员资料进行统计推断

DOI:
10.1111/rssb.12505
复制
发表时间:
2022
期刊:
Journal of the Royal Statistical Society: Series B (Statistical Methodology
影响因子:
--
通讯作者:
Lv, Jinchi
Lv, Jinchi
中科院分区:
--
文献类型:
--
作者:
Fan, Jianqing;Fan, Yingying;Han, Xiao;Lv, Jinchi

文献摘要

参考文献

被引文献

相似文献

网络数据在许多当代大数据应用中普遍存在,其中的共同兴趣是揭示不同节点对之间重要的潜在链接。然而,一个简单的基本问题,即如何精确地量化与识别潜在联系相关的统计不确定性,仍然在很大程度上未被探索。本文在度修正混合成员关系模型的背景下,提出了大型网络成员关系统计推断方法(SIMPLE),其中零假设节点对共享相同的社区成员关系。在更简单的情况下,没有程度的异质性,该模型减少到混合隶属度模型,也提出了一种替代更强大的测试。这两种检验都是基于经验特征向量或其比值的Hotelling型统计量,其渐近协方差矩阵的推导和估计非常具有挑战性。然而,他们的解析表达式被揭开,未知的协方差矩阵的一致估计。在较弱的正则性条件下,分别在原假设和邻接备择假设下,给出了两种形式的SIMPLE检验统计量的精确极限分布。它们分别是卡方分布和非中心卡方分布,其自由度取决于是否校正。我们还解决了重要的问题,估计未知数的社区,并建立相关的检验统计量的渐近性质。通过几个仿真实例和真实的网络应用,证明了我们的新程序在规模和功率方面的优势和实际效用。
Network data are prevalent in many contemporary big data applications in which a common interest is to unveil important latent links between different pairs of nodes. Yet a simple fundamental question of how to precisely quantify the statistical uncertainty associated with the identification of latent links still remains largely unexplored. In this paper, we propose the method of statistical inference on membership profiles in large networks (SIMPLE) in the setting of degree-corrected mixed membership model, where the null hypothesis assumes that the pair of nodes share the same profile of community memberships. In the simpler case of no degree heterogeneity, the model reduces to the mixed membership model for which an alternative more robust test is also proposed. Both tests are of the Hotelling-type statistics based on the rows of empirical eigenvectors or their ratios, whose asymptotic covariance matrices are very challenging to derive and estimate. Nevertheless, their analytical expressions are unveiled and the unknown covariance matrices are consistently estimated. Under some mild regularity conditions, we establish the exact limiting distributions of the two forms of SIMPLE test statistics under the null hypothesis and contiguous alternative hypothesis. They are the chi-square distributions and the noncentral chi-square distributions, respectively, with degrees of freedom depending on whether the degrees are corrected or not. We also address the important issue of estimating the unknown number of communities and establish the asymptotic properties of the associated test statistics. The advantages and practical utility of our new procedures in terms of both size and power are demonstrated through several simulation examples and real network applications.
DOI: --
发表时间: 2017
期刊:
影响因子: --
作者:
Jiashun Jin;Z. Ke
通讯作者: Z. Ke
DOI: 10.1080/07350015.2020.1798241
发表时间: 2017-06
影响因子: 3
作者:
C. Brownlees;Guðmundur Guðmundsson;G. Lugosi
通讯作者: C. Brownlees;Guðmundur Guðmundsson;G. Lugosi