A feature group weighting method for subspace clustering of high-dimensional data

A feature group weighting method for subspace clustering of high-dimensional data
复制标题

DOI:
10.1016/j.patcog.2011.06.004
复制
发表时间:
2012
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
Xiaojun Chen;Yunming Ye;Xiaofei Xu;J. Huang
Xiaojun Chen;Yunming Ye;Xiaofei Xu;J. Huang
中科院分区:
其他
文献类型:
--
作者:
Xiaojun Chen;Yunming Ye;Xiaofei Xu;J. Huang

文献摘要

被引文献

相似文献

本文提出了一种新的方法来加权子空间的特征组和个人的功能聚类高维数据。在该方法中,高维数据的特征被划分为特征组,根据他们的自然特性。两种类型的权重被引入到聚类过程中,以同时识别每个聚类中的特征组和单个特征的重要性。给出了一种新的优化模型来定义优化过程,并提出了一种新的聚类算法FG-k-means来优化优化模型。新算法是对k-means算法的扩展,通过增加两个额外的步骤来自动计算两种类型的子空间权重。提出了一种新的数据生成方法,该方法可以在特征组和单个特征的子空间中生成具有聚类的高维数据。对合成数据和真实数据的实验结果表明,FG-k-means算法明显优于四种k-means类型算法,即,k-means,W-k-means,LAC和EWKM在几乎所有的实验。新算法对高维数据中普遍存在的噪声和缺失值具有较好的鲁棒性。
This paper proposes a new method to weight subspaces in feature groups and individual features for clustering high-dimensional data. In this method, the features of high-dimensional data are divided into feature groups, based on their natural characteristics. Two types of weights are introduced to the clustering process to simultaneously identify the importance of feature groups and individual features in each cluster. A new optimization model is given to define the optimization process and a new clustering algorithm FG-k-means is proposed to optimize the optimization model. The new algorithm is an extension to k-means by adding two additional steps to automatically calculate the two types of subspace weights. A new data generation method is presented to generate high-dimensional data with clusters in subspaces of both feature groups and individual features. Experimental results on synthetic and real-life data have shown that the FG-k-means algorithm significantly outperformed four k-means type algorithms, i.e., k-means, W-k-means, LAC and EWKM in almost all experiments. The new algorithm is robust to noise and missing values which commonly exist in high-dimensional data.