Benchmarking community detection methods on social media data

Benchmarking community detection methods on social media data
复制标题

社交媒体数据上的社区检测方法基准测试

DOI:
--
复制
发表时间:
2013
期刊:
arXiv.org
影响因子:
--
通讯作者:
P. Cunningham
P. Cunningham
中科院分区:
--
文献类型:
--
作者:
Conrad Lee;P. Cunningham

文献摘要

被引文献

相似文献

在经验社会网络数据上对社区检测方法的性能进行基准测试已被确定为改进这些方法的关键。特别是,虽然目前大多数研究都侧重于从大型社交媒体和电信服务中以数字方式提取的数据中检测社区,但对这一研究的大多数评估都是基于小型、手工整理的数据集。我们认为,这两种类型的网络差异如此之大,以至于仅在前者上评估算法,我们对后者的表现知之甚少。为了解决这个问题,我们考虑了基于数字提取的网络数据构建基准时出现的困难,并提出了一种基于任务的策略,我们认为该策略可以解决这些困难。为了证明我们的方案是有效的,我们使用它来执行基于Facebook数据的大量基准测试。基准测试显示,一些最流行的算法无法检测到细粒度的社区结构。
Benchmarking the performance of community detection methods on empirical social network data has been identified as critical for improving these methods. In particular, while most current research focuses on detecting communities in data that has been digitally extracted from large social media and telecommunications services, most evaluation of this research is based on small, hand-curated datasets. We argue that these two types of networks differ so significantly that by evaluating algorithms solely on the former, we know little about how well they perform on the latter. To address this problem, we consider the difficulties that arise in constructing benchmarks based on digitally extracted network data, and propose a task-based strategy which we feel addresses these difficulties. To demonstrate that our scheme is effective, we use it to carry out a substantial benchmark based on Facebook data. The benchmark reveals that some of the most popular algorithms fail to detect fine-grained community structure.