Modeling Over-Dispersion for Network Data Clustering

Modeling Over-Dispersion for Network Data Clustering
复制标题

DOI:
10.1109/icmla.2017.0-180
复制
发表时间:
2017-12
期刊:
2017 16th IEEE International Conference on Machine Learning and Applications (ICMLA)
影响因子:
--
通讯作者:
Lu Wang;D. Zhu;Ming Dong;Yan Li
Lu Wang;D. Zhu;Ming Dong;Yan Li
中科院分区:
其他
文献类型:
--
作者:
Lu Wang;D. Zhu;Ming Dong;Yan Li

文献摘要

相似文献

过分散网络数据挖掘已成为数据科学领域的一个重要研究课题,现实世界中具有不平衡聚类的网络数据量急剧增加,而现有的聚类方法大多是为了发现聚类数目和类间连接模式而设计的,但很少有方法能够发现网络社区和图像片段中普遍存在的不平衡聚类。我们提出了一个通用的概率建模框架,SizeConnectivity,估计过分散的集群大小分布以及来自网络数据的类特定连接模式。我们进行了广泛的合成和真实的-世界实验聚类社交网络数据和图像数据检测网络社区和图像片段。我们的结果表明,我们的SizeConnectivity聚类方法在恢复通过对过度分散进行建模来隐藏网络数据结构。
Over-dispersed network data mining has emerged as a central theme in data science, evident by a sharp increase in the volume of real-world network data with imbalanced clusters.While most of existing clustering methods are designed for discovering the number of clusters and class specific connectivity patterns, few methods are available to uncover the imbalanced clusters,commonly existing in network communities and image segments.In this paper, we propose a generalized probabilistic modeling framework,SizeConnectivity, to estimate over-dispersed cluster size distribution together with class specific connectivity patterns from network data.We performed extensive synthetic and real-world experiments on clustering social network data and image data for detecting network communities and image segments.Our results demonstrate a superior performance of our SizeConnectivity clustering method in recovering the hidden structure of network data via modeling over-dispersion.