Consistent estimation of the number of communities in stochastic block models using cross‐validation

Consistent estimation of the number of communities in stochastic block models using cross‐validation
复制标题

DOI:
10.1002/sta4.426
复制
发表时间:
2021-09
期刊:
影响因子:
1.7
通讯作者:
Jining Qin;Jing Lei
Jining Qin;Jing Lei
中科院分区:
数学4区
文献类型:
--
作者:
Jining Qin;Jing Lei

文献摘要

相似文献

随机块模型(SBM)及其变体构成了研究网络数据的一个重要的概率工具家族。关于随机块模型的块标号和模型参数的估计方法有丰富的文献。这些研究大多需要社区数量K作为输入,使得K的估计成为一个重要的问题。交叉验证是解决这个问题的自然选择,因为它是一种广泛使用的评估模型拟合的通用方法。然而,交叉验证是不一致的,并且倾向于过拟合,除非使用不切实际的分割比。交叉验证与信心(CVC)提出了更好的理论保证,在传统的设置。研究了随机块模型的CVC的性质。我们的理论研究表明,与标准交叉验证不同,CVC可以在适当的条件下始终如一地选择最佳K。我们实现了这种方法,并检查其性能对其他已建立的方法在合成和真实的数据集。
The stochastic block model (SBM) and its variants constitute an important family of probabilistic tools for studying network data. There is a rich literature on methods for estimating block labels and model parameters of stochastic block models. Most of these studies require the number of communities K as an input, making the estimation of K an important problem. Cross‐validation is a natural option for this problem since it is a widely used generic method for evaluating model fitting. However, cross‐validation is known to be inconsistent and prone to overfitting unless impractical split ratios are used. Cross‐validation with confidence (CVC) is proposed with better theoretical guarantees in conventional settings. We study the properties of CVC for stochastic block models. Our theoretical studies show that CVC, unlike the standard cross‐validation, can consistently pick the optimal K under suitable conditions. We implement this method and check its performance against other established methods on both synthetic and real datasets.