Clustering the Google Distance with Eigenvectors and Semidefinite Programming
Clustering the Google Distance with Eigenvectors and Semidefinite Programming
复制标题
使用特征向量和半定规划对 Google 距离进行聚类
DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
T. Zeugmann
中科院分区:
文献类型:
--
作者:
J. Poland;T. Zeugmann
Web mining techniques are becoming increasingly popular and more accurate, as the information body of the World Wide Web grows and reflects a more and more comprehensive picture of the humans’ view of the world. One simple web mining tool is called the Google distance and has been recently suggested by Cilibrasi and Vitányi. It is an information distance between two terms in natural language, and can be derived from the “similarity metric”, which is defined in the context of Kolmogorov complexity. The Google distance can be calculated from just counting how often the terms occur in the web (page counts), e.g. using the Google search engine. In this work, we compare two clustering methods for quickly and fully automatically decomposing a list of terms into semantically related groups: Spectral clustering and clustering by semidefinite programming.