TW-k-Means: Automated Two-Level Variable Weighting Clustering Algorithm for Multiview Data

TW-k-Means: Automated Two-Level Variable Weighting Clustering Algorithm for Multiview Data
复制标题

TW-k-Means:多视图数据的自动两级可变加权聚类算法

DOI:
10.1109/tkde.2011.262
复制
发表时间:
2013-04-01
影响因子:
8.9
通讯作者:
Ye, Yunming
Ye, Yunming
中科院分区:
计算机科学2区
文献类型:
--
作者:
Chen, Xiaojun;Xu, Xiaofei;Ye, Yunming

文献摘要

被引文献

相似文献

本文提出了TW-K均值,这是一种用于多视图数据的自动化的两级变量加权聚类算法,该算法可以同时计算视图和单个变量的权重。在此算法中,将视图权重分配给每个视图,以识别视图的紧凑性,并且在视图中还将可变权重分配给每个变量,以识别变量的重要性。距离函数中都使用了视图权重和可变权重来确定对象的簇。在新算法中,将两个其他步骤添加到迭代K-均值聚类过程中,以自动计算视图权重和可变权重。我们使用两个现实生活数据集研究了TW-K-均值中两种权重的性质,并研究了TW-K-均值的权重与单个可变加权方法的权重之间的差异。实验揭示了TW-K均值中视图权重的收敛性。我们将TW-K-均值与三种现实数据集的五种聚类算法进行了比较,结果表明,TW-K-均值算法在四个评估指数中显着胜过其他五个聚类算法。
This paper proposes TW-k-means, an automated two-level variable weighting clustering algorithm for multiview data, which can simultaneously compute weights for views and individual variables. In this algorithm, a view weight is assigned to each view to identify the compactness of the view and a variable weight is also assigned to each variable in the view to identify the importance of the variable. Both view weights and variable weights are used in the distance function to determine the clusters of objects. In the new algorithm, two additional steps are added to the iterative k-means clustering process to automatically compute the view weights and the variable weights. We used two real-life data sets to investigate the properties of two types of weights in TW-k-means and investigated the difference between the weights of TW-k-means and the weights of the individual variable weighting method. The experiments have revealed the convergence property of the view weights in TW-k-means. We compared TW-k-means with five clustering algorithms on three real-life data sets and the results have shown that the TW-k-means algorithm significantly outperformed the other five clustering algorithms in four evaluation indices.