Web Clustering Based On Tag Set Similarity
Web Clustering Based On Tag Set Similarity
复制标题
DOI:
10.4304/jcp.6.1.59-66
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
Jing-li Zhou;Xuejun Nie;Leihua Qin;Jianfeng Zhu
中科院分区:
文献类型:
--
作者:
Jing-li Zhou;Xuejun Nie;Leihua Qin;Jianfeng Zhu
Tagging is a service that allows users to associate a set of freely determined tags with web content. Clustering web documents with tag sets can eliminate the time-consuming preprocess of word stemming. This paper proposes a novel method to compute the similarity between tag sets and use it as the distance measure to cluster web documents into groups. Major steps in this method include computing a tag similarity matrix with set-based vector space model, smoothing the similarity matrix to obtain a set of linearly independent vectors and compute the tag set similarity based on these vectors. The experimental results show that the proposed tag set similarity measures surpasses other common similarity measures not only in the reliable derivation of clustering results, but also in clustering accuracies and efficiencies.