Effect of Dimensionality Reduction on Different Distance Measures in Document Clustering

Effect of Dimensionality Reduction on Different Distance Measures in Document Clustering
复制标题

文档聚类中降维对不同距离度量的影响

DOI:
--
复制
发表时间:
2011
期刊:
International Conference on Neural Information Processing
影响因子:
--
通讯作者:
T. Honkela
T. Honkela
中科院分区:
--
文献类型:
--
作者:
Mari;Ilkka Kivimäki;Santosh Tirunagari;E. Oja;T. Honkela

文献摘要

被引文献

相似文献

在文档聚类中,语义相似的文档被分组在一起。文档集合的维度往往非常大,数千或数万个术语。因此,由于计算原因,通常在聚类之前降低原始维度。余弦距离被广泛认为是k-均值聚类中度量文档间距离的最佳选择。在本文中,我们对三种降维方法进行了实验,结果表明,在降维到较小的目标维度后,余弦度量的优势不再存在。此外,对于较小的维度,主成分分析降维方法优于奇异值分解方法。我们还展示了L 2归一化如何影响不同的距离度量。这些实验针对三个英语文档集和一个印地语文档集运行。
In document clustering, semantically similar documents are grouped together. The dimensionality of document collections is often very large, thousands or tens of thousands of terms. Thus, it is common to reduce the original dimensionality before clustering for computational reasons. Cosine distance is widely seen as the best choice for measuring the distances between documents in k-means clustering. In this paper, we experiment three dimensionality reduction methods with a selection of distance measures and show that after dimensionality reduction into small target dimensionalities, such as 10 or below, the superiority of cosine measure does not hold anymore. Also, for small dimensionalities, PCA dimensionality reduction method performs better than SVD. We also show how l 2 normalization affects different distance measures. The experiments are run for three document sets in English and one in Hindi.