Bayesian community detection for networks with covariates

Bayesian community detection for networks with covariates
复制标题

具有协变量的网络的贝叶斯社区检测

DOI:
10.48550/arxiv.2203.02090
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
Lizhen Lin
Lizhen Lin
中科院分区:
--
文献类型:
--
作者:
Luyi W. Shen;A. Amini;Nathaniel Josephs;Lizhen Lin

文献摘要

参考文献

被引文献

相似文献

网络数据在各个领域的日益普及以及从中提取有用信息的需求推动了相关模型和算法的快速发展。在网络数据的各种学习任务中,社区检测,即节点簇或“社区”的发现,可以说在科学界受到了最多的关注。在许多现实世界的应用中,网络数据通常带有节点或边协变量形式的附加信息,理想情况下,这些信息应该用于推理。在本文中,我们增加了有限的文献社区检测网络的协变量,提出了一个贝叶斯随机块模型与协变量相关的随机分区先验。在我们的先验下,协变量在指定聚类成员资格的先验分布时被明确表示。我们的模型具有建模的不确定性,包括社区成员的所有参数估计的灵活性。重要的是,与大多数现有方法不同,我们的模型能够通过后验推理来学习社区的数量,而不必假设它是已知的。我们的模型可以应用于稠密和稀疏网络中的社区检测,分类和连续的协变量,我们的MCMC算法是非常有效的,具有良好的混合特性。我们证明了我们的模型优于现有模型的上级性能在一个全面的模拟研究和应用程序的两个真实的数据集。
The increasing prevalence of network data in a vast variety of fields and the need to extract useful information out of them have spurred fast developments in related models and algorithms. Among the various learning tasks with network data, community detection, the discovery of node clusters or"communities,"has arguably received the most attention in the scientific community. In many real-world applications, the network data often come with additional information in the form of node or edge covariates that should ideally be leveraged for inference. In this paper, we add to a limited literature on community detection for networks with covariates by proposing a Bayesian stochastic block model with a covariate-dependent random partition prior. Under our prior, the covariates are explicitly expressed in specifying the prior distribution on the cluster membership. Our model has the flexibility of modeling uncertainties of all the parameter estimates including the community membership. Importantly, and unlike the majority of existing methods, our model has the ability to learn the number of the communities via posterior inference without having to assume it to be known. Our model can be applied to community detection in both dense and sparse networks, with both categorical and continuous covariates, and our MCMC algorithm is very efficient with good mixing properties. We demonstrate the superior performance of our model over existing models in a comprehensive simulation study and an application to two real datasets.
DOI: 10.1214/19-aos1820
发表时间: 2017-09
期刊: The Annals of Statistics
影响因子: --
作者:
E. Kolaczyk;Lizhen Lin;S. Rosenberg;Jie Xu;Jackson Walters
通讯作者: E. Kolaczyk;Lizhen Lin;S. Rosenberg;Jie Xu;Jackson Walters
DOI: 10.1080/01621459.2019.1706541
发表时间: 2016-07
影响因子: 3.7
作者:
Bowei Yan;Purnamrita Sarkar
通讯作者: Bowei Yan;Purnamrita Sarkar
DOI: 10.1016/j.jspi.2010.03.002
发表时间: 2010-10-01
影响因子: 0.9
作者:
Muellner, Peter;Quintana, Fernando
通讯作者: Quintana, Fernando