Clustering the Google Distance with Eigenvectors and Semidefinite Programming

Clustering the Google Distance with Eigenvectors and Semidefinite Programming
复制标题

使用特征向量和半定规划对 Google 距离进行聚类

DOI:
--
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
T. Zeugmann
T. Zeugmann
中科院分区:
--
文献类型:
--
作者:
J. Poland;T. Zeugmann

文献摘要

被引文献

相似文献

随着万维网信息量的增长和越来越全面地反映人类对世界的看法,Web挖掘技术正变得越来越流行和准确。一个简单的网络挖掘工具被称为Google距离,最近由Cilibrasi和Vitányi提出。它是自然语言中两个术语之间的信息距离,并且可以从“相似性度量”中导出,该相似性度量在Kolmogorov复杂度的上下文中定义。Google距离可以通过计算这些词在网络中出现的频率(页面计数)来计算,例如使用Google搜索引擎。在这项工作中,我们比较了两种聚类方法,快速,完全自动地分解成语义相关的组的术语列表:谱聚类和聚类半定规划。
Web mining techniques are becoming increasingly popular and more accurate, as the information body of the World Wide Web grows and reflects a more and more comprehensive picture of the humans’ view of the world. One simple web mining tool is called the Google distance and has been recently suggested by Cilibrasi and Vitányi. It is an information distance between two terms in natural language, and can be derived from the “similarity metric”, which is defined in the context of Kolmogorov complexity. The Google distance can be calculated from just counting how often the terms occur in the web (page counts), e.g. using the Google search engine. In this work, we compare two clustering methods for quickly and fully automatically decomposing a list of terms into semantically related groups: Spectral clustering and clustering by semidefinite programming.