Modeling Over-Dispersion for Network Data Clustering
Modeling Over-Dispersion for Network Data Clustering
复制标题
DOI:
10.1109/icmla.2017.0-180
复制
发表时间:
2017-12
期刊:
影响因子:
--
通讯作者:
Lu Wang;D. Zhu;Ming Dong;Yan Li
中科院分区:
文献类型:
--
作者:
Lu Wang;D. Zhu;Ming Dong;Yan Li
Over-dispersed network data mining has emerged as a central theme in data science, evident by a sharp increase in the volume of real-world network data with imbalanced clusters.While most of existing clustering methods are designed for discovering the number of clusters and class specific connectivity patterns, few methods are available to uncover the imbalanced clusters,commonly existing in network communities and image segments.In this paper, we propose a generalized probabilistic modeling framework,SizeConnectivity, to estimate over-dispersed cluster size distribution together with class specific connectivity patterns from network data.We performed extensive synthetic and real-world experiments on clustering social network data and image data for detecting network communities and image segments.Our results demonstrate a superior performance of our SizeConnectivity clustering method in recovering the hidden structure of network data via modeling over-dispersion.