Effects of maximum flow algorithm on identifying web community

Effects of maximum flow algorithm on identifying web community
复制标题

最大流量算法对网络社区识别的影响

DOI:
10.1145/584931.584941
复制
发表时间:
2002
期刊:
--
影响因子:
--
通讯作者:
M. Kitsuregawa
M. Kitsuregawa
中科院分区:
--
文献类型:
--
作者:
Noriko Imafuji;M. Kitsuregawa

文献摘要

被引文献

相似文献

在本文中,我们描述了使用最大流量算法从网络中提取网络社区的效果。网络社区是一组具有共同主题的网页。由于网络可以被认为是由分别代表网页和超链接的节点和边组成的图,因此迄今为止已经提出了各种图论方法来从网络图中提取网络社区。利用最大流量算法寻找网络社区的方法是普林斯顿NEC研究所在两年前提出的。然而,通过这种方法得出的网络社区的属性却很少为人所知。为了检验该方法的效果,我们随机选取了30个主题,并利用2000年爬取的日本网络档案进行了实验。通过这些实验,我们发现该方法既有优点也有缺点。我们将描述一些有效使用此方法的策略。此外,通过使用相同的主题,我们研究了另一种基于完全二分图的方法。我们比较了通过这些方法获得的网络社区并分析了这些特征。
In this paper, we describe the effects of using maximum flow algorithm on extracting web community from the web. A web community is a set of web pages having a common topic. Since the web can be recognized as a graph that consists of nodes and edges that represent web pages and hyperlinks respectively, so far various graph theoretical approaches have been proposed to extract web communities from the web graph. The method of finding a web community using maximum flow algorithm was proposed by NEC Research Institute in Princeton two years ago. However the properties of web communities derived by this method have been seldom known. To examine the effects of this method, we selected 30 topics randomly and experimented using Japanese web archives crawled in 2000. Through these experiments, it became clear that the method has both advantages and disadvantages. We will describe some strategies to use this method effectively. Moreover, by using same topics, we examined another method that is based on complete bipartite graphs. We compared the web communities obtained by those methods and analyzed those characteristics.