Web Clustering Based On Tag Set Similarity

Web Clustering Based On Tag Set Similarity
复制标题

DOI:
10.4304/jcp.6.1.59-66
复制
发表时间:
2011
期刊:
J. Comput.
影响因子:
--
通讯作者:
Jing-li Zhou;Xuejun Nie;Leihua Qin;Jianfeng Zhu
Jing-li Zhou;Xuejun Nie;Leihua Qin;Jianfeng Zhu
中科院分区:
其他
文献类型:
--
作者:
Jing-li Zhou;Xuejun Nie;Leihua Qin;Jianfeng Zhu

文献摘要

被引文献

相似文献

标签是一种允许用户将一组自由确定的标签与web内容相关联的服务。使用标记集聚类web文档可以消除耗时的词词干提取预处理。本文提出了一种计算标签集之间相似度的新方法,并将其作为网络文档聚类的距离度量。该方法的主要步骤是利用基于集合的向量空间模型计算标签相似度矩阵,对相似度矩阵进行平滑处理,得到一组线性无关的向量,并基于这些向量计算标签集相似度。实验结果表明,所提出的标签集相似度度量不仅在聚类结果的可靠推导方面优于其他常用的相似度度量,而且在聚类精度和效率方面也优于其他常用的相似度度量。
Tagging is a service that allows users to associate a set of freely determined tags with web content. Clustering web documents with tag sets can eliminate the time-consuming preprocess of word stemming. This paper proposes a novel method to compute the similarity between tag sets and use it as the distance measure to cluster web documents into groups. Major steps in this method include computing a tag similarity matrix with set-based vector space model, smoothing the similarity matrix to obtain a set of linearly independent vectors and compute the tag set similarity based on these vectors. The experimental results show that the proposed tag set similarity measures surpasses other common similarity measures not only in the reliable derivation of clustering results, but also in clustering accuracies and efficiencies.