Stratified feature sampling method for ensemble clustering of high dimensional data

Stratified feature sampling method for ensemble clustering of high dimensional data
复制标题

高维数据集成聚类的分层特征采样方法

DOI:
10.1016/j.patcog.2015.05.006
复制
发表时间:
2015-11-01
影响因子:
8
通讯作者:
Huang, Joshua Z.
Huang, Joshua Z.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jing, Liping;Tian, Kuang;Huang, Joshua Z.

文献摘要

被引文献

相似文献

高维数据具有数千个特征,这对现有的聚类算法提出了很大的挑战。稀疏性、噪声和特征的相关性是此类数据的共同特征。另一个常见的现象是,这种高维数据中的聚类往往存在于不同的子空间中。包围式聚类是一种提高高维数据聚类鲁棒性、稳定性和准确性的重要技术。本文提出了一种在高维数据集成聚类中生成子空间分量数据集的分层抽样方法。在这种方法中,我们首先将高维数据的特征聚类成几个特征组,称为特征层,而不是随机地为每个组件数据集抽取一个特征子集。采用分层抽样的方法,从每个特征层中随机抽取一些特征,并将不同特征层中抽取的特征进行合并,生成一个组件数据集。通过这种方式,组件数据集具有原始数据集中聚类结构的更好表示。与综合数据分析中的随机抽样和随机投影方法相比,分层抽样的成分聚类在不牺牲聚类多样性的前提下,提高了平均聚类精度。我们进行了一系列的实验,8个真实的世界的数据集,从微阵列,文本和图像域使用三个子空间分量数据生成方法和四个共识函数的集成聚类方法进行评估。实验结果一致表明,分层抽样方法产生的最好的集成聚类结果在所有的数据集。分层抽样的集成聚类也优于其他三种集成聚类方法,这些方法从原始数据的整个空间生成组件聚类。(C)2015爱思唯尔有限公司版权所有。
High dimensional data with thousands of features present a big challenge to current clustering algorithms. Sparsity, noise and correlation of features are common characteristics of such data. Another common phenomenon is that clusters in such high dimensional data often exist in different subspaces. Ensemble clustering is emerging as a prominent technique for improving robustness, stability and accuracy of high dimensional data clustering. In this paper, we propose a stratified sampling method for generating subspace component data sets in ensemble clustering of high dimensional data. Instead of randomly sampling a subset of features for each component data set, in this method we first cluster the features of high dimensional data into a few feature groups called feature strata. Using stratified sampling, we randomly sample some features from each feature stratum and merge the sampled features from different feature strata to generate a component data set. In this way, the component data sets have better representations of the clustering structure in the original data set. Comparing with random sampling and random projection methods in synthetic data analysis, the component clustering by stratified sampling has demonstrated that the average clustering accuracy was increased without sacrificing clustering diversity. We carried out a series of experiments on eight real world data sets from microarray, text and image domains to evaluate ensemble clustering methods using three subspace component data generation methods and four consensus functions. The experimental results consistently showed that the stratified sampling method produced the best ensemble clustering results in all data sets. The ensemble clustering with stratified sampling also outperformed three other ensemble clustering methods which generate component clusters from the entire space of the original data. (C) 2015 Elsevier Ltd. All rights reserved.