A Method for Pinpoint Clustering of Web Pages with Pseudo-Clique Search
A Method for Pinpoint Clustering of Web Pages with Pseudo-Clique Search
复制标题
DOI:
10.1007/11605126_4
复制
发表时间:
2005-05
期刊:
影响因子:
--
通讯作者:
M. Haraguchi;Yoshiaki Okubo
中科院分区:
文献类型:
--
作者:
M. Haraguchi;Yoshiaki Okubo
This paper presents a method forPinpoint Clusteringof web pages. We try to find useful clusters of web pages which are significant in the sense that their contents are similar to ones of higher-ranked pages. Since we are usually careless of lower-ranked pages, they are unconditionally discarded even if their contents are similar to some pages with high ranks. Such hidden pages together with significant higher-ranked pages are extracted as a cluster. As the result, our clusters can provide new valuable information for users.In order to obtain such clusters, we first extract semantic correlations among terms by applyingSingular Value Decomposition(SVD) to the term-document matrix generated from a corpus. Based on the correlations, we can evaluate potential similarities among web pages to be clustered. The set of web pages is represented as a weighted graphGbased on the similarities and their ranks. Our clusters can be found aspseudo-cliquesinG. An algorithm for finding Top-Nweighted pseudo-cliques is presented. Our experimental result shows that a quite valuable cluster can be actually extracted according to our method.We also discuss an idea for improvement on meanings of clusters. With the help ofFormal Concept Analysis, our clusters, called FC-based clusters, can be provided with clear meanings. Our preliminary experimentation shows that the extended method would be a promising approach to finding meaningful clusters.