Description-Driven Community Detection

Description-Driven Community Detection
复制标题

DOI:
10.1145/2517088
复制
发表时间:
2014-04
期刊:
ACM Trans. Intell. Syst. Technol.
影响因子:
--
通讯作者:
Simon Pool;F. Bonchi;M. Leeuwen
Simon Pool;F. Bonchi;M. Leeuwen
中科院分区:
其他
文献类型:
--
作者:
Simon Pool;F. Bonchi;M. Leeuwen

文献摘要

被引文献

相似文献

传统的社区检测方法,如物理学家、社会学家和最近的计算机科学家所研究的,旨在简单地划分社交网络图。然而,随着在线社交网站的出现,更丰富的数据已经变得可用:除了链接信息之外,网络中的每个用户都被注释了附加信息,例如人口统计学,购物行为或兴趣。因此,在这方面,必须开发能够利用所有现有信息的采矿方法。在社区检测的情况下,这意味着找到良好的社区(社交图中的一组节点),这些社区与用户信息(节点属性)方面的良好描述相关联。与我们的模型相关联的良好描述使它们能够被领域专家理解,从而在现实世界的应用程序中更有用。另一个由现实世界的应用程序决定的需求是开发可以使用任何特定领域背景知识的方法。在社区检测的情况下,背景知识可以是在特定应用中寻找的社区的模糊描述,或者是一些原型节点(例如,过去的好客户),这代表了分析师正在寻找的东西(类似用户的社区)。为了实现这一目标,在这篇文章中,我们定义和研究的问题,找到一个不同的一组有凝聚力的社区与简洁的描述。我们提出了一个有效的算法,交替两个阶段:爬山阶段生产(可能重叠)社区,和描述归纳阶段,使用技术从监督模式集挖掘。我们的框架具有很好的功能,能够从任何给定的节点描述或种子集开始构建描述良好的内聚社区,这使得它非常灵活,易于应用于现实世界的应用程序。我们的实验评估证实,所提出的方法发现凝聚力的社区与简洁的描述,在现实和大型的在线社交网络,如美味,Flickr和LastFM。
Traditional approaches to community detection, as studied by physicists, sociologists, and more recently computer scientists, aim at simply partitioning the social network graph. However, with the advent of online social networking sites, richer data has become available: beyond the link information, each user in the network is annotated with additional information, for example, demographics, shopping behavior, or interests. In this context, it is therefore important to develop mining methods which can take advantage of all available information. In the case of community detection, this means finding good communities (a set of nodes cohesive in the social graph) which are associated with good descriptions in terms of user information (node attributes). Having good descriptions associated to our models make them understandable by domain experts and thus more useful in real-world applications. Another requirement dictated by real-world applications, is to develop methods that can use, when available, any domain-specific background knowledge. In the case of community detection the background knowledge could be a vague description of the communities sought in a specific application, or some prototypical nodes (e.g., good customers in the past), that represent what the analyst is looking for (a community of similar users). Towards this goal, in this article, we define and study the problem of finding a diverse set of cohesive communities with concise descriptions. We propose an effective algorithm that alternates between two phases: a hill-climbing phase producing (possibly overlapping) communities, and a description induction phase which uses techniques from supervised pattern set mining. Our framework has the nice feature of being able to build well-described cohesive communities starting from any given description or seed set of nodes, which makes it very flexible and easily applicable in real-world applications. Our experimental evaluation confirms that the proposed method discovers cohesive communities with concise descriptions in realistic and large online social networks such as Delicious, Flickr, and LastFM.