A Method for Pinpoint Clustering of Web Pages with Pseudo-Clique Search

A Method for Pinpoint Clustering of Web Pages with Pseudo-Clique Search
复制标题

DOI:
10.1007/11605126_4
复制
发表时间:
2005-05
期刊:
--
影响因子:
--
通讯作者:
M. Haraguchi;Yoshiaki Okubo
M. Haraguchi;Yoshiaki Okubo
中科院分区:
其他
文献类型:
--
作者:
M. Haraguchi;Yoshiaki Okubo

文献摘要

被引文献

相似文献

本文提出了一种网页精确定位的方法。我们试图找到有用的集群的网页,这是显着的意义上说,他们的内容是类似的排名较高的网页。由于我们通常对排名较低的页面不太在意,因此即使它们的内容与一些排名较高的页面相似,它们也会被无条件地丢弃。这样的隐藏页面与显著的较高排名的页面一起被提取为集群。为了得到这样的聚类,我们首先对语料库中的术语-文档矩阵进行奇异值分解(SVD),提取术语之间的语义相关性。基于这些相关性,我们可以评估待聚类网页之间的潜在相似性。这组网页根据相似性及其排名表示为加权图G。我们的集群可以被发现是相互独立的。提出了一种求Top-N加权伪团的算法。实验结果表明,该方法可以提取出有价值的聚类,并提出了改进聚类意义的思路。在形式概念分析的帮助下,我们的聚类,称为基于FC的聚类,可以提供明确的含义。我们的初步实验表明,扩展的方法将是一个很有前途的方法来找到有意义的集群。
This paper presents a method forPinpoint Clusteringof web pages. We try to find useful clusters of web pages which are significant in the sense that their contents are similar to ones of higher-ranked pages. Since we are usually careless of lower-ranked pages, they are unconditionally discarded even if their contents are similar to some pages with high ranks. Such hidden pages together with significant higher-ranked pages are extracted as a cluster. As the result, our clusters can provide new valuable information for users.In order to obtain such clusters, we first extract semantic correlations among terms by applyingSingular Value Decomposition(SVD) to the term-document matrix generated from a corpus. Based on the correlations, we can evaluate potential similarities among web pages to be clustered. The set of web pages is represented as a weighted graphGbased on the similarities and their ranks. Our clusters can be found aspseudo-cliquesinG. An algorithm for finding Top-Nweighted pseudo-cliques is presented. Our experimental result shows that a quite valuable cluster can be actually extracted according to our method.We also discuss an idea for improvement on meanings of clusters. With the help ofFormal Concept Analysis, our clusters, called FC-based clusters, can be provided with clear meanings. Our preliminary experimentation shows that the extended method would be a promising approach to finding meaningful clusters.