Document Clustering in Correlation Similarity Measure Space

Document Clustering in Correlation Similarity Measure Space
复制标题

相关相似性测度空间中的文档聚类

DOI:
10.1109/tkde.2011.49
复制
发表时间:
2012-06-01
影响因子:
8.9
通讯作者:
Xiang, Yong
Xiang, Yong
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhang, Taiping;Tang, Yuan Yan;Xiang, Yong

文献摘要

被引文献

相似文献

本文提出了一种新的谱聚类方法--相关保持索引(CPI),它是在相关相似性度量空间中进行的。在这个框架中,文件被投影到一个低维的语义空间中,在该空间中,本地补丁中的文件之间的相关性被最大化,而这些补丁之外的文件之间的相关性同时被最小化。由于文档空间的内在几何结构往往嵌入在文档之间的相似性中,相关性作为相似性度量比欧氏距离更适合于检测文档空间的内在几何结构。因此,提出的CPI方法可以有效地发现嵌入在高维文档空间的内在结构。在不同的数据集上进行了大量的实验,并与现有的文档聚类方法进行了比较,证明了新方法的有效性。
This paper presents a new spectral clustering method called correlation preserving indexing (CPI), which is performed in the correlation similarity measure space. In this framework, the documents are projected into a low-dimensional semantic space in which the correlations between the documents in the local patches are maximized while the correlations between the documents outside these patches are minimized simultaneously. Since the intrinsic geometrical structure of the document space is often embedded in the similarities between the documents, correlation as a similarity measure is more suitable for detecting the intrinsic geometrical structure of the document space than euclidean distance. Consequently, the proposed CPI method can effectively discover the intrinsic structures embedded in high-dimensional document space. The effectiveness of the new method is demonstrated by extensive experiments conducted on various data sets and by comparison with existing document clustering methods.