Clustering by Pattern Similarity

Clustering by Pattern Similarity
复制标题

按模式相似度进行聚类

DOI:
--
复制
发表时间:
2008
期刊:
Journal of Computational Science and Technology
影响因子:
--
通讯作者:
J. Pei
J. Pei
中科院分区:
--
文献类型:
--
作者:
Haixun Wang;J. Pei

文献摘要

被引文献

相似文献

聚类的任务是在一组对象中识别相似对象的类别。不同的聚类模型对相似性的定义不同。然而,在大多数这些模型中,相似性的概念通常是基于诸如曼哈顿距离、欧几里得距离或其他Lp距离等度量。换句话说,相似的对象必须至少在一组维度上具有相近的值。在本文中,我们探索了一种更一般的相似性类型。在我们提出的pCluster模型下,如果两个对象在维度子集上表现出一致的模式,则它们是相似的。新的相似性概念模型具有广泛的应用前景。例如,在DNA微阵列分析中,两个基因的表达水平可能在一系列环境刺激下同步上升和下降。虽然它们的表达水平的大小可能不接近,但它们表现出的模式可能非常相似。这些基因簇的发现对于揭示基因调控网络中的重要联系至关重要。电子商务应用,如协同过滤,也可以从这个新模型中受益,因为它不仅可以捕捉到某些领先指标的价值接近度,还可以捕捉到客户表现出的模式(购买、浏览等)的接近度。除了提出新的相似度模型外,本文还提出了一种有效的聚类检测算法,并在多个真实数据集和合成数据集上进行了测试。
The task of clustering is to identify classes of similar objects among a set of objects. The definition of similarity varies from one clustering model to another. However, in most of these models the concept of similarity is often based on such metrics as Manhattan distance, Euclidean distance or other Lp distances. In other words, similar objects must have close values in at least a set of dimensions. In this paper, we explore a more general type of similarity. Under the pCluster model we proposed, two objects are similar if they exhibit a coherent pattern on a subset of dimensions. The new similarity concept models a wide range of applications. For instance, in DNA microarray analysis, the expression levels of two genes may rise and fall synchronously in response to a set of environmental stimuli. Although the magnitude of their expression levels may not be close, the patterns they exhibit can be very much alike. Discovery of such clusters of genes is essential in revealing significant connections in gene regulatory networks. E-commerce applications, such as collaborative filtering, can also benefit from the new model, because it is able to capture not only the closeness of values of certain leading indicators but also the closeness of (purchasing, browsing, etc.) patterns exhibited by the customers. In addition to the novel similarity model, this paper also introduces an effective and efficient algorithm to detect such clusters, and we perform tests on several real and synthetic data sets to show its performance.